NVIDIA Vera Rubin NVL72 debuts in MLPerf Inference v6.1 with up to 3.7x throughput over GB300 NVL72
NVIDIA submitted its first MLPerf Inference v6.1 preview results for Vera Rubin NVL72, showing up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL. The numbers signal a leap for rack-scale inference on MoE models.