Summary
- Nebius has acquired Inferize and folded its team into the Token Factory inference platform.
- Inferize focuses on shortening model cold starts that can leave allocated GPU capacity idle.
- The deal extends Nebius’s investment beyond hardware towards the software determining production inference cost and utilisation.
Nebius has acquired inference optimisation startup Inferize, adding software designed to cut model startup delays as AI infrastructure providers look beyond acquiring more GPUs and towards extracting more useful work from the hardware they already operate.
Amsterdam-headquartered Nebius said Inferize’s engineering team and technology will join Token Factory, its managed service for production inference workloads. Financial terms were not disclosed.
Inferize works on the delay between allocating computing capacity and having a model loaded and ready to handle requests. Those cold starts can appear when capacity expands to meet demand or when a workload switches between models, leaving expensive accelerators allocated but temporarily unable to serve users.
The cost becomes more pronounced as AI moves from demonstrations towards applications with unpredictable traffic. Training runs consume large amounts of compute over a defined period, whereas production inference has to respond when users or software agents arrive, maintain latency targets, and absorb sudden changes in demand.
Utilisation becomes part of the cloud price
Infrastructure providers can preserve performance by keeping spare GPU capacity ready for demand peaks, although idle hardware carries the same capital and power costs while producing no billable work. Nebius describes the effect as an “idle GPU tax”, a vendor phrase for an operational problem shared across production AI systems.
Reducing startup time allows capacity to follow demand more closely. If model instances can be prepared faster, a provider needs less idle hardware sitting in reserve and can increase the proportion of its fleet performing useful inference at any one time.
That changes how infrastructure costs should be compared. The advertised price of a GPU hour captures only part of the economics because throughput, model loading, utilisation, scheduling, memory, and networking all influence how many requests the same hardware can process.
A cheaper accelerator that spends more time waiting can produce a higher effective cost per useful unit of work, while software improvements can make existing hardware more competitive without changing the underlying chip. Inference optimisation has consequently become part of the contest between AI cloud providers rather than an engineering detail left entirely to customers.
Inferize was founded in January 2026 and, according to Nebius, produced its first working prototype within three months. Its engineers will now work across Token Factory, beginning with integration of the startup’s technology into the managed platform.
Nebius is buying further up the stack
The acquisition follows other moves into inference software. Nebius agreed to acquire Eigen AI in May, adding optimisation work across models, kernels, and systems, while it also brought the core Clarifai team into the business and licensed technology related to inference and compute orchestration.
Those transactions extend the company’s strategy beyond supplying accelerator capacity. Nebius still needs large fleets of GPUs and the data centres, networking, and power behind them, but its commercial performance also depends on how efficiently customers can use that infrastructure.
Token Factory packages some of that engineering as a managed service, allowing developers to consume production endpoints without operating the full serving stack themselves. The model competes with cloud platforms that increasingly bundle hardware, model hosting, optimisation, monitoring, and scaling into a single service.
European AI infrastructure providers face an additional scale problem because the largest US cloud groups can spread software investment across much larger customer bases. Specialised providers therefore need advantages in areas such as model performance, access to capacity, price, sovereignty, or engineering efficiency if they are to avoid competing solely on rented GPU supply.
Inferize gives Nebius another piece of software aimed at improving that equation, although the company has not published production benchmarks showing the effect once the technology is integrated across Token Factory. The acquisition case rests for now on the value of cutting wasted accelerator time rather than independently demonstrated savings at fleet scale.
As inference becomes a larger share of AI spending, those margins will attract more attention. The industry spent the first part of the current AI cycle racing to secure chips; providers now have an equally strong incentive to ensure that the chips already installed are not sitting idle while a model loads.










