On August 24, 2026, NVIDIA published preview results showing that its Vera Rubin NVL72 rack-scale system achieves up to 30x higher AI-factory throughput per megawatt than the GB300 NVL72 on the SemiAnalysis AgentX benchmark. The AgentX workload replays production-style coding agent sessions — long-context prefill, KV-cache reuse, tool-call gaps, and dynamic concurrency — using the AIPerf client. At a sustained 160 tokens per second per user on the DeepSeek V4-Pro model, Vera Rubin NVL72 delivers the 30x efficiency gain while maintaining the same interactive serving target.
The same benchmark shows GB300 NVL72 extending its multi-generational advantage over H200 NVL8. On DeepSeek V4 Pro 1.6T, GB300 NVL72 delivers up to 15x higher throughput per megawatt and up to 10x lower cost per million tokens. On the larger Kimi K3 2.8T mixture-of-experts model, the advantage widens to 80x throughput per megawatt, with the interactivity frontier reaching roughly 215 tokens per second per user — well beyond H200 NVL8's operating range.
What's new
- Vera Rubin NVL72: Up to 30x higher AI-factory throughput per megawatt vs. GB300 NVL72 on AgentX (DeepSeek V4-Pro, 160 tok/s/user). Results measured by NVIDIA; pending SemiAnalysis review.
- GB300 NVL72: Up to 15x throughput per megawatt vs. H200 NVL8 on DeepSeek V4 Pro 1.6T; up to 10x lower token cost.
- GB300 NVL72 on Kimi K3 2.8T: Up to 80x throughput per megawatt vs. H200 NVL8; interactivity frontier ~215 tok/s/user.
- Benchmark: SemiAnalysis AgentX (InferenceX suite) — open-source, replays recorded Claude Code sessions with interleaved reasoning and tool use via AIPerf client.
- Key metrics: Tokens per megawatt reported against E2E normalized interactivity, standard interactivity, E2E latency, and TTFT.
Why it matters
Agentic AI has shifted inference from single-turn chat to multi-step workflows that reason, invoke tools, coordinate subagents, and carry growing context across turns. OpenRouter's State of AI report found average prompt tokens per request grew roughly fourfold across 100 trillion tokens of real-world usage, and single agentic requests consume 15x the tokens of ordinary chat. This makes tokens-per-megawatt the decisive economic metric for AI factories. Vera Rubin NVL72's 30x gain over Blackwell on a realistic agentic workload signals that rack-scale co-design — spanning GPUs, CPUs, NVLink fabric, and the Dynamo serving stack — can convert a fixed power budget into dramatically more interactive agentic capacity.
Our take
The 30x figure is a NVIDIA-measured preview pending independent SemiAnalysis review at a specific operating point. Still, the progression — H200 to GB300 (15x) to Vera Rubin (30x over GB300) — shows consistent architecture-level leverage on agentic workloads that static benchmarks miss. The Vera CPU's role in tool execution and KV-cache offload, plus NVLink 6 and Dynamo's session-aware routing, suggest the bottleneck has moved from raw GPU FLOPs to system-level data movement and orchestration.
Sources
- NVIDIA Technical Blog: Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt
- NVIDIA Newsroom: NVIDIA Vera Rubin Opens Agentic AI Frontier
- NVIDIA Technical Blog: Vera CPU Sets a New Standard for Agentic Workloads in AI Factories
- SiliconANGLE: Nvidia Vera Rubin: Inside the agentic AI factory that rewrites the CPU playbook