NVIDIA Demonstrates Qwen3.8-Flash-Next 176B on GB300 NVL72 for Agentic Coding
NVIDIA validated Alibaba's Qwen3.8-Flash-Next 176B MoE model on its GB300 NVL72 rack, achieving over 16K tokens per second per GPU. The hybrid GDN/QSA architecture keeps memory and compute bounded for million-token agentic coding workloads.