NVIDIA Serves Alibaba's Qwen3.8-2.4T-A95B on GB300 NVL72 With Day-0 FP8 Performance
NVIDIA published Day-0 FP8 inference results for Alibaba's 2.4T-parameter Qwen3.8-2.4T-A95B on the GB300 NVL72 rack, hitting over 4,000 tokens per second per GPU. The result demonstrates that rack-scale NVLink can absorb MoE all-to-all traffic at frontier model scale.