When Arrcus introduced its AI Inference Network Fabric (AINF) ahead of MWC 26, we saw an AI-focused pitch on distributed inference. Standing back, the more interesting story is where AINF sits in the stack. Model selection is a solved, widely deployed problem — an application caller like OpenRouter picks among hosted models. Worker and pod scheduling inside a site is mature, through frameworks such as NVIDIA Dynamo and vLLM. What has been missing is the layer between them: a policy-aware fabric that decides which site serves a request under real-time latency, power, and sovereignty constraints.
AINF is that layer. It is a Kubernetes-native fabric, deployed next to the inference runtime at each site, that steers every request across a footprint of edge, data center, and cloud. A router at each site evaluates capacity and load, model and adapter availability, conversation affinity for KV-cache reuse, geo-fencing, and service tier in real time, per the AINF announcement. Arrcus cites third-party research suggesting over 60% lower time-to-first-token and 40% lower end-to-end latency; we treat those as directional rather than apples-to-apples.
As a transparent overlay, an AINF Router accepts OpenAI-compatible requests and routes at L7 over mTLS, feeding load and KV-cache telemetry to a distributed control plane while a central Orchestrator distributes policy and aggregates fabric-wide telemetry. The diagram below shows a request landing on any of five sites across three regions — a regional site when latency matters, a sovereign site for data residency, or an overflow site when local capacity is constrained.

Sovereignty is a first-class feature: requests can be pinned to jurisdictions, and the Orchestrator’s telemetry supports audit reporting down to the model and request. Peering routers use SRv6/MPLS traffic-engineered paths across the WAN; policy-aware inference routers and ToR switches run on xPU and NVIDIA Spectrum/BF3 and Broadcom silicon with a range of ODM hardware. We think this is the right wedge: as AI infrastructure matures, guaranteeing throughput, time-to-first-token, geofencing, and sovereign capacity stops being a model-selection problem and becomes a routing-fabric problem — and few vendors sell that layer for the carrier footprint. 650 Group researches several of the technologies in this post, including AI Networking and Data Center, and more.