On August 24, 2026, NVIDIA introduced Scale-In as the fifth pillar of its AI networking infrastructure, joining Scale-Up, Scale-Out, Scale-Across, and Context Memory. Scale-In targets the north-south access path that connects users, agents, applications, data sources, and storage systems to the accelerated compute inside an AI factory.

The architecture is powered by the NVIDIA BlueField-4 DPU, the DOCA software platform, and Spectrum-X Ethernet. BlueField-4 provides dedicated, host-independent acceleration for policy enforcement, storage access, security, and telemetry at up to 800 Gb/s, offloading these infrastructure services from host CPUs so they do not become bottlenecks as AI compute scales.

What's new

  • Scale-In pillar: Evolves north-south networks into a coordinated infrastructure domain for the AI factory, handling access, security, data movement, and operations.
  • BlueField-4 DPU: Integrates a 64-core NVIDIA Grace CPU, inline acceleration engines, LPDDR5X memory, PCIe Gen6 host connection, and an 800 Gb/s network interface. Compared with BlueField-3, it delivers 4x more memory bandwidth and 2x more network bandwidth.
  • DOCA microservices: Containerized services run directly on BlueField-4 for networking, security, storage, and telemetry. DOCA Flow programs hardware packet pipelines, DOCA PCC handles programmable congestion control, DOCA Telemetry exposes health metrics, and DOCA Platform Framework manages provisioning and updates.
  • Spectrum-X Ethernet: Provides the high-performance fabric across the Scale-In access path, addressing load-balancing conflicts and congestion at scale while isolating concurrent traffic for predictable performance.
  • BlueField Astra: Extends trusted control from Scale-In into the east-west Scale-Out fabric, giving service providers a unified, host-independent control point for provisioning, tenant isolation, and network policy across both domains.

How Scale-In fits the AI factory

NVIDIA defines five complementary infrastructure pillars. Scale-Up uses NVLink to unite GPUs as a coherent accelerator. Scale-Out connects servers across the factory with Spectrum-X Ethernet and Quantum InfiniBand. Scale-Across links distributed factories via Spectrum-XGS Ethernet. Context Memory (CMX), built on the STX modular foundation, provides pod-level shared KV-cache storage for faster inference. Scale-In now addresses the north-south access, security, data movement, and operations surrounding the compute domain.

In the Vera Rubin NVL72 platform, ConnectX-9 SuperNICs carry tenant workload traffic over the Scale-Out network while BlueField-4 runs the infrastructure services that connect, secure, and manage each server. This separation keeps infrastructure processing outside the tenant host, preventing security, data access, and operations from consuming host CPU resources.

Key use cases for agentic AI factories

  • Isolated AI factory virtual private clouds: DOCA Host-Based Networking accelerates north-south Layer 3 routing and multi-tenant isolation on BlueField-4. DOCA Flow programs traffic classification and access-control rules, while DOCA-accelerated Open vSwitch applies policy on east-west interfaces. BlueField Astra extends the same VPC policies across Scale-In and Scale-Out.
  • High-performance storage access: Scale-In provides accelerated access to AI factory storage for training, inference, and analytics, as well as enterprise AI data systems for retrieval and multimodal indexing. BlueField-4 serves as the infrastructure processor for Scale-In and, separately, as the data and storage processor for CMX.
  • Runtime threat detection and telemetry: Host-independent processing enables continuous security monitoring and fleet-wide observability without relying on host CPU cycles.

Why it matters

Agentic AI workloads turn infrastructure into part of the inference pipeline. Each request can trigger many model calls, tool calls, memory lookups, policy checks, and storage accesses before a final answer is produced. As context windows grow to millions of tokens and agents run continuously, the north-south path must move, protect, and retrieve data at line rate without introducing latency or consuming host CPU cycles needed for agent execution. Scale-In addresses this by making infrastructure services programmable, accelerated, and isolated from tenant workloads.

Our take

NVIDIA is formalizing a networking layer that cloud providers have been building piecemeal. By co-designing BlueField-4, DOCA, and Spectrum-X with Vera Rubin, the company aims to remove the host CPU from the infrastructure data path entirely — a prerequisite for gigascale AI factories where every megawatt must translate into tokens per second. The real test will be adoption beyond NVIDIA reference platforms and whether DOCA's microservice model attracts a broad ISV ecosystem.

Series: 1. NVIDIA Vera BlueField-4 STX Benchmarks Show Up to 3.7x Faster Storage Primitives for AI-Native Workloads · 2. NVIDIA Draws a Security Line for AI Agents: Harness Guides, Runtime Decides · 3. Amazon Science releases SOP-Bench to test AI agents on real business procedures · 4. NVIDIA Introduces Scale-In as Fifth AI Networking Pillar with BlueField-4 DPU · Agent Infrastructure

Sources