On September 16, 2026, NVIDIA published its first MLPerf Inference v6.1 preview results for the Vera Rubin NVL72 platform. The submission shows up to 3.7x higher throughput than the previous-generation GB300 NVL72 on the Qwen3-VL vision-language benchmark and up to 2.5x on the DeepSeek-R1 reasoning model, using vLLM with the NVIDIA Dynamo framework and TensorRT-LLM respectively.

Related: Earlier coverage: NVIDIA Vera Rubin NVL72 Delivers 30x Throughput-per-Watt Gain Over Blackwell on AgentX Benchmark.

The results come from MLPerf Inference v6.1 Closed Division entries 6.1-0106 and 6.1-0074, retrieved from MLCommons on the same day. NVIDIA emphasizes that these are early preview numbers and that performance will improve with continued software optimization.

Confirmed

  • Vera Rubin NVL72 preview submission on DeepSeek-R1 and Qwen3-VL-235B-A22B benchmarks in MLPerf Inference v6.1.
  • Up to 3.7x throughput improvement over GB300 NVL72 on Qwen3-VL across offline, server, and interactive scenarios (vLLM + Dynamo).
  • Up to 2.5x throughput improvement over GB300 NVL72 on DeepSeek-R1 (TensorRT-LLM).
  • GB300 NVL72 scaled from 1 rack (72 GPUs) to 4 racks (288 GPUs) with 99% scaling efficiency on DeepSeek-R1 offline (entries 6.1-0073 and 6.1-0074).
  • Software optimizations in v6.1 delivered up to 1.6x higher performance on Qwen3-VL over v6.0 results on the same GB300 NVL72 hardware.
  • Post-submission gains reported on GPT-OSS-120B and DLRMv3, not yet verified by MLCommons.
  • Nebius also submitted Vera Rubin NVL72 preview results.
  • 19 partners participated in this round, eight on multi-node Blackwell NVL72 systems.
  • Jetson AGX Thor results submitted on the new Edge-Agentic benchmark with Qwen3.6-27B using TensorRT Edge-LLM.

Unknown

  • Independent replication of Vera Rubin NVL72 preview numbers; MLCommons has not yet verified the submission.
  • General availability timeline and pricing for Vera Rubin NVL72 systems.
  • Full power and TCO comparison at production scale; CoreWeave measured 10x tokens-per-megawatt on engineering samples, SemiAnalysis estimates 5.4x perf-per-MW and 5x perf-per-dollar over GB200 NVL72, both from early bring-up.
  • Software maturity of the Rubin stack (CUDA 13.4, SM_107) and kernel optimization runway.
  • Competitive submissions from Google TPU v7 and AMD MI455X UALoE72, expected later in 2026.
  • MLPerf Endpoints benchmark for agentic workloads is upcoming; current agentic claim (30x over GB300 NVL72 on SemiAnalysis AgentX) is from preview testing only.

Our take

NVIDIA's preview submission signals that Vera Rubin's architectural changes — sixth-gen NVLink, NVFP4 precision, disaggregated serving, and expert parallelism — are translating into measurable throughput gains on mixture-of-experts models that dominate agentic workloads. The 99% scaling efficiency on GB300 NVL72 also demonstrates that the rack-scale interconnect and software orchestration are maturing together. However, the Vera Rubin numbers remain preview-grade: they come from engineering samples, lack independent MLCommons verification, and will shift as the software stack matures. Buyers evaluating 2027 infrastructure should treat the 3.7x figure as a directional upper bound, not a committed SLA.

Series: 1. Nvidia Commits Up to $3 Billion to Lancium, Powering the Stargate AI Campus in Texas · 2. NVIDIA Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x throughput over GB300 NVL72 · NVIDIA AI Factory

Sources