DeepMind After Hassabis: What Actually Changes Under Koray Kavukcuoglu?
Series end: what Google confirmed under Kavukcuoglu’s SVP seat — and what is still analysis or unknown. Reporting map + Confirmed / Analysis / Unknown.
114 results for “Anthropic”
Series end: what Google confirmed under Kavukcuoglu’s SVP seat — and what is still analysis or unknown. Reporting map + Confirmed / Analysis / Unknown.
NVIDIA validated Alibaba's Qwen3.8-Flash-Next 176B MoE model on its GB300 NVL72 rack, achieving over 16K tokens per second per GPU. The hybrid GDN/QSA architecture keeps memory and compute bounded for million-token agentic coding workloads.
NVIDIA unveiled NVLink Fusion, letting hyperscalers plug custom XPUs into its 72‑XPU NVLink domain and MGX rack architecture. The program cuts non‑differentiated engineering so silicon teams can ship custom accelerators faster while reusing proven power, cooling, and cluster software.
NVIDIA declared Scale-In the fifth pillar of AI networking infrastructure, powered by BlueField-4 DPUs, DOCA software, and Spectrum-X Ethernet. The architecture offloads security, storage, and telemetry from host CPUs at up to 800 Gb/s to keep pace with agentic AI factories.
NVIDIA detailed how Spectrum-X Ethernet uses hardware-accelerated adaptive routing, congestion control, and plane load balancing to solve traditional Ethernet's failures at giga-scale AI training. The platform delivers 1.6x performance over off-the-shelf Ethernet, near-perfect multi-tenant isolation, and 2.68 ms failover versus 1.08 seconds for standard Ethernet.
NVIDIA took a minority stake in data-center site developer Cloverleaf Infrastructure to accelerate AI factory construction across the U.S. The partnership integrates NVIDIA's DSX platform early in site design to maximize compute output per megawatt.
NVIDIA launched DSX MaxLPS, a full-stack platform that dynamically reallocates stranded power across GPU racks to fit up to 40% more compute in the same facility envelope. The system combines real-time power steering, workload-aware profiles, and 45°C liquid cooling to boost performance per watt by 1.3–1.5x on Blackwell and Vera Rubin systems.
NVIDIA published a practical observability framework mapping telemetry tools to every layer of DGX and HGX AI factory infrastructure. The guide helps teams catch gray failures like InfiniBand bit-error drift before they waste multi-hour training runs.
NVIDIA published a tutorial showing how KAI Scheduler and vCluster let three teams share one L40S GPU while each gets an isolated Kubernetes control plane. The pattern scales to thousands of nodes and gives platform teams a way to consolidate GPU hardware without sacrificing tenant autonomy.
NVIDIA published Day-0 FP8 inference results for Alibaba's 2.4T-parameter Qwen3.8-2.4T-A95B on the GB300 NVL72 rack, hitting over 4,000 tokens per second per GPU. The result demonstrates that rack-scale NVLink can absorb MoE all-to-all traffic at frontier model scale.
Nvidia has formally opened the cuFile APIs and the vertical storage stack beneath them, moving the interface that lets GPUs read and write storage. Opening cuFile is a strategic departure; this layer has historically lived inside CUDA.
Nvidia has agreed to invest up to $3 billion in Lancium, the Texas-based power infrastructure developer behind the first operational site of the Stargate. Power has replaced silicon as the scarcest resource in frontier AI development.