Summary
- Nebius will raise selected pay-as-you-go GPU prices by 17% to 21% from 1 October, its second increase in three months.
- Some CPU-only instances and memory products are also becoming more expensive, while longer commitments continue to receive discounts.
- Higher infrastructure prices complicate the economics of moving AI workloads from experimentation into persistent production.
Nebius is raising prices across parts of its AI cloud estate from 1 October, providing another indication that the rapid build-out of GPU infrastructure has not yet translated into steadily cheaper on-demand compute.
The Amsterdam-based cloud provider will increase pay-as-you-go rates for selected Nvidia GPUs by between 17% and 21%, while some CPU-only instances will rise by 25% and prices for certain memory products will increase by around 41%.
It is the company’s second price increase in three months, arriving as demand for accelerator capacity remains strong across model training, fine-tuning, and inference. Nebius has also been expanding its European infrastructure footprint, including high-density Nvidia deployments in Finland and the UK, while competing with specialist AI cloud providers and hyperscalers for workloads requiring large volumes of advanced chips.
Its current public pricing lists on-demand rates ranging from $3.85 per GPU-hour for Nvidia H100 infrastructure to $7.15 for B200 systems, with H200 capacity priced at $4.50. Lower rates remain available for pre-emptible capacity, while customers reserving large clusters over longer periods can receive commitment discounts of up to 35%.
AI budgets are turning into infrastructure budgets
The increase arrives at an awkward point in enterprise AI adoption because the economics change materially once an application moves from an occasional pilot to an always-on service. A proof of concept can tolerate bursts of expensive compute when usage is limited to a small technical team, whereas inference embedded in customer service, software engineering, search, content processing, or internal operations accumulates cost continuously.
Infrastructure procurement consequently starts to resemble conventional capacity planning. Businesses have to decide how much compute to reserve, how much to leave on demand, which models justify premium accelerators, and whether workloads can tolerate pre-emptible instances or cheaper hardware. Utilisation, latency targets, data residency, contractual commitments, and portability between clouds all affect the eventual unit cost.
Nebius’s decision also illustrates the difference between ordinary cloud computing and the current AI-infrastructure market. Conventional cloud economics have been shaped by mature hardware supply chains and years of competition around standardised compute, whereas the most sought-after AI systems depend on a narrower pool of accelerators, high-speed networking, power availability, and data-centre capacity capable of supporting dense racks.
Prices can therefore remain firm even as newer hardware becomes more efficient. More capable accelerators may reduce the amount of hardware needed for a particular workload, but that advantage is diluted when demand grows faster than usable capacity or organisations consume the additional performance with larger models, longer contexts, richer multimodal workloads, and more persistent inference.
Commitments cut cost while reducing flexibility
Nebius is maintaining lower prices for customers willing to commit to capacity, which is useful to a provider funding expensive infrastructure and potentially attractive to organisations with predictable demand. The trade-off is familiar from the broader cloud market: a lower unit rate is exchanged for a stronger forecast about future usage.
That calculation becomes harder in AI because application architecture remains fluid. A workload designed around one model family or accelerator generation can look very different six months later if a smaller model becomes adequate, an inference optimisation reduces demand, or a competing provider changes the economics. Committed capacity can protect against price rises and shortages, but it can also preserve assumptions that stop being optimal.
The pricing move follows a period of rapid commercial expansion for Nebius, which has been signing large customer agreements while increasing infrastructure capacity. Rival specialist GPU clouds have also pointed to firm contract pricing, suggesting that persistent demand is not confined to a single supplier.
European organisations building AI products will feel the broader effect through procurement rather than model choice alone. FinOps teams increasingly need to treat GPUs as a scarce and variable infrastructure input, engineering teams face stronger pressure to measure utilisation and inference efficiency, and finance departments have to distinguish headline model costs from the cloud, storage, networking, and reserved capacity required to keep those models available.
AI infrastructure continues to expand quickly, although a second Nebius increase in three months shows why greater installed capacity does not automatically produce lower spot prices. Until accelerator supply, power, networking, and demand settle into a more mature balance, the cost of keeping an AI system running may remain considerably less predictable than the cost of demonstrating that it works.












