Summary
- Gartner forecasts AI-optimised IaaS spending will rise 96.4% to $42.3 billion during 2026.
- Inference is expected to account for $23.3 billion, overtaking the $19 billion spent on model training.
- The forecast suggests AI infrastructure costs are shifting towards continuous use inside operational applications and agentic workflows.
Gartner expects companies to spend more cloud infrastructure money running AI models than training them during 2026, as the economics of artificial intelligence begin shifting from building models towards paying continuously for the applications and agents using them.
Worldwide spending on AI-optimised infrastructure as a service is forecast to reach $42.3 billion this year, up 96.4% from $21.5 billion in 2025. Gartner expects the market to reach $66.1 billion during 2027, meaning specialised AI infrastructure is growing substantially faster than the broader IaaS market even as cloud spending remains strong overall.
The composition of that expenditure is more revealing than the headline growth rate. Gartner expects inference workloads to consume $23.3 billion during 2026, compared with $19 billion for training, giving inference 55% of AI-optimised IaaS spending. Its share is forecast to rise again to 59% in 2027.
Training attracts attention because large model runs require enormous clusters of accelerators for concentrated periods, but a model that enters production creates a different cost profile. Every customer query, document processed, recommendation generated, code task completed, or agentic workflow executed requires inference, turning successful adoption into an ongoing infrastructure expense rather than a one-off development bill.
The distinction becomes more important as companies move AI from experimental interfaces into business processes. A pilot used by a few hundred employees may create manageable consumption, whereas a system embedded in customer service, software development, search, document processing, fraud detection, or internal operations can generate requests throughout the working day and sometimes continuously.
Agentic systems can add considerably more compute because one user request may lead to several model calls. An agent can plan a task, retrieve information, call tools, inspect results, revise its approach, and continue until the workflow is complete, meaning infrastructure consumption is not necessarily visible from the number of people using the application.
Gartner identifies that multi-step execution as one of the forces making inference the dominant consumption model. As fine-tuned and domain-specific models move into customer-facing and operational systems, computing demand becomes persistent because the models have to execute whenever the underlying business process runs.
Cloud budgeting consequently becomes part of the AI business case in a different way. Organisations can calculate the cost of an initial training or development project relatively clearly, but production economics depend on usage volumes, model choice, context length, response patterns, orchestration, and the number of times a workflow invokes a model before completing its task.
Those variables can make apparently successful adoption financially awkward. If an AI application becomes widely used while each interaction remains expensive, higher utilisation can increase the cloud bill faster than the process savings it was intended to produce. Smaller models, selective routing, caching, or redesigned workflows can reduce consumption without abandoning the application.
AI cost management is consequently beginning to resemble cloud financial operations more than a conventional software-licensing exercise. Computing becomes a metered input, and optimisation involves technical decisions about which model handles which task, where inference runs, how many calls are made, and whether every stage genuinely requires the most capable available system.
The trend strengthens the commercial position of infrastructure providers because production workloads can recur for years across accelerators, networking, storage, databases, and supporting cloud services. An agent embedded in an operational system can generate consumption every time the organisation performs the relevant process.
Customers therefore have a stronger reason to separate model capability from infrastructure efficiency. The highest-performing model on a benchmark may not produce the best economics for a high-volume business task if a smaller or more specialised alternative delivers adequate performance at materially lower inference cost.
Europe adds another layer because cloud economics interact with requirements around data location, resilience, regulatory governance, and supplier concentration. Some workloads will remain on major public-cloud platforms, while others may justify private, sovereign, or locally operated infrastructure when utilisation and control requirements make the additional engineering worthwhile.
None of those alternatives makes inference free. Whether computing is bought through a hyperscaler, reserved on specialist infrastructure, or operated internally, persistent model use consumes processors, power, and engineering capacity. Moving the workload changes where the bill appears and who controls the infrastructure rather than eliminating it.
Gartner’s wider forecast puts the scale in context. Total IaaS spending is expected to rise from $222.2 billion in 2025 to $287.3 billion this year and $359.9 billion in 2027, so AI-optimised infrastructure remains a minority of the overall cloud market while growing far more quickly.
As inference becomes the larger portion of AI infrastructure spending, adoption metrics will need to mature beyond how many employees have access to a tool or how many models a company has deployed. The more useful numbers will include how much each workflow costs to run, how frequently it is invoked, and whether the operating benefit still exceeds the computing bill once AI becomes an everyday production system.












