Summary
- Gartner forecasts that inference costs per agentic AI workflow will increase more than fivefold through 2028 despite continuing improvements in the unit economics of AI models.
- The research argues that multistep agents consume substantially more inference than simple chatbot interactions because they reason, route tasks, call tools, and repeatedly evaluate intermediate results.
- Controlling costs will require organisations and software suppliers to route work between models deliberately rather than defaulting every task to the most capable autonomous system.
Falling prices for artificial-intelligence models may do considerably less to reduce enterprise AI bills than software suppliers and customers expect, with Gartner forecasting that inference costs for an individual agentic workflow will increase more than fivefold through 2028. The research argues that improvements in the price of individual tokens are being overtaken by the amount of computation required as AI moves from answering questions towards carrying out longer sequences of work.
The distinction cuts across one of the more reassuring assumptions surrounding enterprise AI economics. Competition between model providers, improvements in hardware, and more efficient software have pushed the cost of many individual AI interactions down, but an autonomous agent does substantially more work than a conventional chatbot before producing a result. It may interpret an objective, plan several steps, call external tools, retrieve information, reconsider earlier actions, pass work to another model, and repeat parts of the process when an intermediate answer appears inadequate.
Gartner describes the resulting effect as an inference paradox: the unit cost of intelligence improves while the total amount of intelligence consumed by an application expands faster. Its forecast suggests that routing a task through an agentic reasoning model already creates inference costs at least five times those of a basic chatbot interaction, with more complicated workflows capable of widening that difference further. The fivefold forecast through 2028 is therefore about the cost of completing a workflow rather than the published price of a model token.
As companies move AI into operational systems rather than using it mainly for occasional writing or information retrieval, the volume of inference becomes a larger part of infrastructure expenditure. Gartner separately forecasts that global spending on AI-optimised infrastructure as a service will reach about $42.3 billion in 2026, up 96.4% from 2025, with inference spending expected to exceed training expenditure this year. Inference is forecast to account for $23.3 billion of that infrastructure spending, compared with $19 billion for training, as more organisations run AI continuously inside customer-facing and operational applications.
Agents multiply the work behind one request
A simple chatbot interaction can often be understood as a relatively contained exchange: a user provides an instruction, the model processes some context, and a response is generated. Agentic systems add orchestration around that model, allowing software to decide what action should happen next and to keep working until a broader objective has been completed. The resulting user experience may appear simpler because somebody can ask for an outcome rather than specify every step, but the computing activity behind that request becomes substantially more elaborate.
An enterprise purchasing agent, for example, might receive an instruction to identify suitable suppliers rather than merely summarise a document. Completing the job could require it to query internal data, search approved external sources, compare products, analyse contractual conditions, check procurement rules, calculate costs, request another model to review its reasoning, and prepare a recommendation. Each stage creates additional model calls and tokens, while higher-stakes decisions can require more capable reasoning models, larger contexts, or repeated verification before the system is allowed to proceed.
The economics are consequently different from software whose marginal cost is close to zero once it has been deployed. Every additional model call consumes compute, and agentic systems can create thousands or millions of those calls once they are embedded across a workforce or customer base. An interaction that looks inexpensive during a pilot can become material when it is invoked continuously across finance, customer service, software engineering, procurement, logistics, or other high-volume processes.
Gartner’s forecast does not mean every agent will inevitably become five times more expensive. It is an aggregate prediction about increasingly capable workflows, and organisations can influence their own costs through architecture, model choice, task design, caching, and limits on autonomy. Published token prices nevertheless provide a poor proxy for the eventual cost of an AI-enabled business process when the number and type of model interactions are themselves changing.
Model routing becomes part of software economics
One response is to avoid sending every problem to the most powerful model available. Many agentic workflows contain tasks of very different difficulty, from extracting a field or classifying an email to making a multi-variable judgement that requires more capable reasoning. If software can route routine operations to cheaper models while reserving expensive systems for the few steps that need them, the cost of the overall workflow can be reduced without necessarily reducing its usefulness.
That creates a more complicated application architecture than simply connecting an enterprise system to one foundation model. Suppliers may need to manage several models, monitor their relative performance, determine when a task should be escalated, and keep those routing rules current as vendors change prices and capabilities. Organisations also have to decide when additional inference produces enough accuracy, reliability, or commercial value to justify the expense rather than allowing an agent to continue reasoning indefinitely.
The trade-off extends into governance because the cheapest routing decision is not always the appropriate one. A low-cost model may be adequate for summarising an internal memo but unsuitable for interpreting a regulated transaction, evaluating a cybersecurity alert, or determining whether a customer should receive a particular financial product. Cost optimisation therefore has to sit alongside controls governing risk, model quality, data access, human review, and the consequences of an incorrect action.
Software pricing may have to adapt as well. Enterprise applications have traditionally been sold through combinations of per-user subscriptions, licences, and usage charges, whereas agentic products can generate substantial variable infrastructure costs depending on what users ask them to do. Suppliers promising broad autonomous functionality for a predictable subscription price have to absorb that variability somewhere, which can place pressure on margins if customers make heavier use of expensive reasoning workflows than pricing models anticipated.
AI infrastructure shifts from training to operation
The broader spending data suggests that this is becoming an operational infrastructure question rather than a concern confined to model developers. Gartner expects 55% of AI-optimised infrastructure-as-a-service spending to support inference in 2026 and 59% in 2027, reflecting the transition from building models towards running them repeatedly inside production systems. Training remains enormously compute-intensive, but it happens comparatively intermittently; a deployed AI service can generate inference demand every time a customer, employee, software system, or automated agent invokes it.
For European organisations, that shift lands alongside concerns over cloud expenditure, data location, energy consumption, and dependence on a relatively small group of model and infrastructure providers. Agentic applications can deepen those dependencies because a workflow designed around a particular model’s reasoning behaviour or tool interface may be harder to move than a simple chatbot connection. Cost control is therefore partly an architecture decision about how easily models can be substituted rather than merely a procurement negotiation over token prices.
The economics also complicate productivity calculations. An AI agent that costs substantially more than a chatbot can still be worthwhile if it automates enough useful work, reduces errors, shortens a process, or allows employees to handle more complex cases. Conversely, a cheap model can destroy value if an organisation automates a low-value process badly and then runs it millions of times. Measuring AI expenditure against the outcome of an entire workflow, rather than against the price of an individual model call, becomes increasingly important as autonomy expands.
Gartner’s forecast undermines the assumption that model efficiency alone will solve application economics. Token prices can continue falling while enterprise AI expenditure rises because the software built on top of those models is asking them to perform more reasoning, take more actions, and remain active for longer. As agentic systems move from demonstrations into production, organisations will have to decide how much intelligence each task actually requires and build software capable of enforcing that choice rather than assuming cheaper models will take care of the bill themselves.












