Summary
- Equinix Inference Exchange combines Nvidia enterprise architectures, Together AI’s inference platform, and Equinix infrastructure and connectivity.
- Together AI will support more than 200 open-source models through both shared multitenant and dedicated single-tenant deployments.
- The service targets latency, model portability, and data-residency requirements and is due to become available from the first quarter of 2027.
As enterprise AI moves into production, the infrastructure question is becoming less about whether organisations can access a capable model and more about where each request should be processed, because latency, data location, cost, and governance can make centralising every inference workload impractical.
Equinix is trying to turn that placement decision into an infrastructure service through Inference Exchange, a collaboration with Nvidia and Together AI that will distribute model-serving capacity across its data-centre estate. The service is due to become available from the first quarter of 2027.
The proposition combines Nvidia’s validated enterprise reference architectures with Together AI’s inference platform, which supports more than 200 open-source models, while Equinix provides the data-cententre infrastructure, power, cooling, operations, and private connectivity through its existing Fabric network.
Together AI will support both multitenant deployments, where infrastructure is shared for efficiency, and dedicated single-tenant environments for organisations requiring isolated capacity. The resulting architecture is intended to connect inference workloads with corporate data, clouds, networks, and other AI providers rather than treating a model endpoint as a self-contained service.
That becomes important once AI applications depend on the rest of the enterprise technology estate. A system answering customer questions, processing operational information, or assisting employees may need to reach databases, identity systems, software platforms, and model providers spread across several environments, while the organisation also has to decide where sensitive information can be processed.
Inference becomes a placement decision
Much of the first enterprise AI investment cycle concentrated on models and accelerators, but inference produces different infrastructure economics from training. A trained model can spend months or years answering requests from applications and users spread across different locations, making response time, connectivity, and per-request cost recurring operational concerns.
Equinix is consequently pitching metropolitan or regional inference as an alternative to routing every request back to a distant central cluster. Bringing processing nearer users or data can reduce network distance, although distributing inference across more locations also creates additional infrastructure that has to be monitored, secured, and capacity-planned.
The company is also targeting organisations that want greater freedom to move between proprietary and open models. Together AI’s catalogue provides another route to run inference without committing every application permanently to one model provider, while Equinix’s interconnection layer is intended to keep access to surrounding cloud and enterprise systems consistent.
Whether that flexibility survives production deployment will depend on more than model choice. Companies can become dependent on the infrastructure, networking, evaluation systems, data pipelines, and proprietary integrations surrounding a model even when the underlying model itself can be changed.
Nevertheless, portability is becoming more commercially relevant as organisations discover that the most capable model is not necessarily the most economical one for every workload. High-volume applications can make relatively small differences in inference cost significant, while regulated or sensitive processes may justify infrastructure controls that would be excessive for a generic productivity assistant.
Data residency reaches the inference layer
Equinix identifies sovereign AI as another intended use case, allowing organisations to select locations intended to support data-residency and sovereignty requirements. The terminology spans a wide range of political and technical objectives, but the underlying operating question is straightforward: sending corporate information to a model means deciding where that processing takes place and which suppliers control the path.
An organisation can keep customer or operational data inside a particular jurisdiction yet weaken that location strategy if the information is routinely transmitted to an inference service elsewhere without equivalent controls. The location of model execution therefore becomes another component of data architecture rather than a decision belonging solely to AI development teams.
Distributed inference creates a trade-off between localisation and operational simplicity. Running services in more places can improve latency and support residency requirements, while simultaneously creating more infrastructure to patch, monitor, secure, and govern. Equinix’s commercial opportunity lies in absorbing some of that work into a managed infrastructure layer.
The company enters the market with more than 280 data centres across 77 metropolitan areas and says eight of the ten largest AI model providers and nine of the ten largest AI cloud providers already deploy infrastructure within its ecosystem. That gives Equinix a dense set of connections around which to build the service, although it does not establish whether enterprises will prefer the package to inference supplied directly by cloud or model providers.
Nvidia’s involvement also shows how the chipmaker is extending its influence beyond accelerator sales. Reference architectures define how compute, networking, storage, and software should be assembled, allowing Nvidia to shape the systems around its hardware even where another company owns the facility and a third operates the inference platform.
For Equinix, the strategy is consistent with an attempt to sit between enterprises and a fragmented infrastructure market rather than compete as another general-purpose cloud. If organisations genuinely use several models, clouds, AI providers, and private environments, the connections between them become valuable infrastructure in their own right.
Inference Exchange still has to demonstrate its performance, pricing, regional footprint, and customer adoption once it launches in 2027. Its architecture nevertheless reflects a more mature phase of enterprise AI in which organisations need to decide not only which model works, but where it should run, what it should connect to, how much each answer costs, and which jurisdiction governs the information travelling through it.












