On September 16, 2026, NVIDIA published its first MLPerf Inference v6.1 preview results for the Vera Rubin NVL72 platform. The submission shows up to 3.7x higher throughput than the previous-generation GB300 NVL72 on the Qwen3-VL vision-language benchmark and up to 2.5x on the DeepSeek-R1 reasoning model, using vLLM with the NVIDIA Dynamo framework and TensorRT-LLM respectively.
The results come from MLPerf Inference v6.1 Closed Division entries 6.1-0106 and 6.1-0074, retrieved from MLCommons on the same day. NVIDIA emphasizes that these are early preview numbers and that performance will improve with continued software optimization.
Confirmed
- Vera Rubin NVL72 preview submission on DeepSeek-R1 and Qwen3-VL-235B-A22B benchmarks in MLPerf Inference v6.1.
- Up to 3.7x throughput improvement over GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios (vLLM + Dynamo).
- Up to 2.5x throughput improvement over GB300 NVL72 on DeepSeek-R1 (TensorRT-LLM).
- GB300 NVL72 scaled from 1 rack (72 GPUs) to 4 racks (288 GPUs) with 99% scaling efficiency on DeepSeek-R1 offline (entries 6.1-0073 and 6.1-0074).
- Software optimizations in v6.1 delivered up to 1.6x higher performance on Qwen3-VL over v6.0 results on the same GB300 NVL72 hardware.
- Post-submission gains reported on GPT-OSS-120B and DLRMv3, not yet verified by MLCommons.
- Nebius also submitted Vera Rubin NVL72 preview results.
- 19 partners participated in this round, eight on multi-node Blackwell NVL72 systems.
- Jetson AGX Thor results submitted on the new Edge-Agentic benchmark with Qwen3.6-27B using TensorRT Edge-LLM.
Unknown
- Independent replication of Vera Rubin NVL72 preview numbers; MLCommons has not yet verified the submission.
- General availability timeline and pricing for Vera Rubin NVL72 systems.
- Full power and TCO comparison at production scale; CoreWeave measured 10x tokens-per-megawatt on engineering samples, SemiAnalysis estimates 5.4x perf-per-MW and 5x perf-per-dollar over GB200 NVL72, both from early bring-up.
- Software maturity of the Rubin stack (CUDA 13.4, SM_107) and kernel optimization runway.
- Competitive submissions from Google TPU v7 and AMD MI455X UALoE72, expected later in 2026.
- MLPerf Endpoints benchmark for agentic workloads is upcoming; current agentic claim (30x over GB300 NVL72 on SemiAnalysis AgentX) is from preview testing only.
Our take
NVIDIA's preview submission signals that Vera Rubin's architectural changes — sixth-gen NVLink, NVFP4 precision, disaggregated serving, and expert parallelism — are translating into measurable throughput gains on mixture-of-experts models that dominate agentic workloads. The 99% scaling efficiency on GB300 NVL72 also demonstrates that the rack-scale interconnect and software orchestration are maturing together. However, the Vera Rubin numbers remain preview-grade: they come from engineering samples, lack independent MLCommons verification, and will shift as the software stack matures. Buyers evaluating 2027 infrastructure should treat the 3.7x figure as a directional upper bound, not a committed SLA.
Series: 1. Nvidia Commits Up to $3 Billion to Lancium, Powering the Stargate AI Campus in Texas · 2. NVIDIA Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x throughput over GB300 NVL72 · NVIDIA AI Factory
Sources
- NVIDIA Blog: Vera Rubin NVL72 Delivers Leading Performance in MLPerf Inference v6.1 Debut
- NVIDIA Developer Blog: Platform Delivers Lowest Token Cost Enabled by Extreme Co-Design
- CoreWeave Blog: First-Ever Measured Vera Rubin NVL72 Silicon Performance Stats
- SemiAnalysis InferenceX: Vera Rubin NVL72 vs GB200 NVL72 Inference TCO & Architecture Analysis