On a sweltering August evening in Silicon Valley, as air conditioning loads spiked, Silicon Valley Power sent a demand signal to an AI factory to reduce its power consumption. Emerald AI’s Conductor platform — a grid-orchestration tool from an NVIDIA partner, running on NVIDIA’s Eos facility — automatically dropped the site’s draw from four megawatts to three in under a minute while keeping every high-priority inference job running. The utility has since sent more than 200 such signals, and NVIDIA says the response worked every time.

That Santa Clara run is not a shipped NVIDIA DSX Flex install. NVIDIA frames it as production proof that AI factories can act as dispatchable load — the pattern DSX Flex is meant to generalize. Separately, at the AI Infra Summit, NVIDIA VP Ian Buck put factory efficiency at the center of his keynote, and cloud provider Lambda released first validation numbers for DSX MaxLPS (also covered in NVIDIA’s summit companion post): on a five-rack, 19-node HGX B200 cluster, running nodes at 85% of full power delivered about 24% more cluster-wide token throughput (roughly 4 million to 5 million tokens per second) and a 23% improvement in performance per watt versus a 16-node full-power baseline.

For Brocker’s power-and-pace thread, this is the infrastructure half of the same constraint discussed in AI labs’ pace-frontier week: tokens per megawatt, not peak FLOPS alone.

Confirmed

  • NVIDIA DSX is positioned as a full-stack AI-factory platform spanning MaxLPS (dynamic power allocation), Flex (grid-signal response), OS (lifecycle), Sim (pre-deployment simulation), and Reference Designs.
  • Lambda DSX MaxLPS validation (HGX B200): +24% token throughput and +23% performance per watt at an 85% power policy versus 16 nodes at full power, per Lambda/NVIDIA figures released with the AI Infra Summit coverage.
  • Emerald AI Conductor on NVIDIA Eos: participant in Silicon Valley Power’s Flexible Load Interconnect Program; automated demand response 4 MW → 3 MW in under a minute; 200+ utility signals with no high-priority job interruption, per NVIDIA. NVIDIA presents this as commercial-scale proof for the Flex pattern, not as a dedicated DSX Flex product deployment.
  • First dedicated DSX Flex commercial deployment (NVIDIA): a 96 MW Vera Rubin AI factory at NVIDIA’s AI Factory Research Center in Manassas, Virginia.
  • NVIDIA projection: DSX MaxLPS can enable up to 40% more GPU capacity for Vera Rubin NVL72 factories within the same megawatt budget in “suitable deployment environments.”
  • 800V DC: NVIDIA says DSX reference designs incorporate 800V DC architecture aimed at denser racks and fewer conversion stages, aligned with Vera Rubin NVL72 timing (NVIDIA cites 2027 availability for that generation’s power path).
  • Summit companion claims (NVIDIA AI Infra Summit post): On agentic / long-context demos, NVIDIA cites Groq 3 LPX on Vera Rubin at up to 35× token throughput per megawatt vs GB200 NVL72 for 2T+ parameter long-context setups, and SemiAnalysis AgentX figures of up to 30× throughput per megawatt (DeepSeek V4 Pro) vs GB300 NVL72 — all vendor- or chart-attributed, not independently audited here.

Unknown

  • Independent replication of Lambda’s MaxLPS results on other hardware mixes and workload mixes.
  • Pricing and licensing for DSX software components (MaxLPS, Flex, OS, Sim).
  • General-availability timeline for DSX Flex beyond the Manassas plan.
  • How widely utilities will offer Flexible Load Interconnect–style programs that pay or prioritize dispatchable AI load.
  • 800V DC OEM/ODM readiness and facility retrofit cost outside NVIDIA’s reference designs.
  • Whether the 40% Vera Rubin capacity uplift holds outside NVIDIA’s “suitable environments” caveat.
  • Broader stack claims in companion NVIDIA developer posts (multi-generation tokens-per-MW multipliers, cuLitho, training energy papers) are out of scope for this piece and remain vendor-reported until independently checked.
  • Groq 3 LPX and SemiAnalysis AgentX multipliers from the summit wrap — workload-specific (e.g. Qwen / DeepSeek configs cited by NVIDIA); independent replication and pricing/SLA for Vera Rubin NVL72 remain open.

Our take

NVIDIA is reframing AI infrastructure from peak FLOPS to tokens per megawatt — a metric that folds silicon, provisioning, and grid contracts into one scoreboard. The Eos + Emerald run shows grid flexibility can work at commercial scale without killing priority inference; DSX Flex is the productization bet, still ahead of its first dedicated site. Economics hinge on utilities that value dispatchable load. Without programs like Silicon Valley Power’s, Flex stays a niche. The 40% GPU-capacity claim for Vera Rubin is a projection with an explicit suitability caveat — treat it as roadmap math, not a sizing guarantee.

Sources