On September 15, 2026, Google DeepMind announced two new models — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking — designed to make voice interactions with AI feel more natural, fluid, and intelligent. Both models are rolling out starting today through the Gemini API, Google AI Studio, Google Workspace, Search Live, and the Gemini app.

Gemini 3.8 Live targets scale and cost efficiency, combining conversational intelligence with real-time visual grounding and background tool execution. Gemini 3.8 Live Extended Thinking is built for high-complexity tasks, delivering increased intelligence and multi-step reasoning while keeping the conversation flowing through early verbal cues and live progress narration.

Model Speech to Speech Index
Gemini 3.8 Live Extended Thinking (High) 82.6%
GPT-Live-1 Astra (Medium) 81.5%
Grok Voice Think Fast 2.0 (High) 81.3%
Gemini 3.8 Live 76.0%
Gemini 3.1 Flash Live (High) 71.5%
Gemini 3.1 Flash Live (Minimal) 63.9%

Artificial Analysis Speech to Speech Index — figures from Google DeepMind’s announcement chart; higher is better. Independent replication pending. Methodology: DeepMind evals note.

Confirmed

  • Availability: 3.8 Live is rolling out starting today for developers in the Gemini API and Google AI Studio; for enterprises in private preview in Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience; for everyone in Search Live. 3.8 Live Extended Thinking is rolling out starting today for developers in the Gemini API and Google AI Studio; for enterprises in private preview in Gemini Enterprise and coming soon to Gemini Enterprise for Customer Experience and Google Workspace business customers; for everyone in Gemini Live and for Google AI Pro and Ultra subscribers in Workspace (Docs) and all Google AI subscribers in Gmail and Keep.
  • Core capabilities: Both models process visual inputs in near real-time, automatically detect and transition between 97 supported languages mid-conversation, and execute tools and API calls in the background while continuing the dialogue.
  • Extended Thinking differentiation: The Extended Thinking variant reasons and speaks simultaneously, using acknowledgments like “Let me check that…” and live progress narration to walk users through multi-step background tasks without interrupting the conversational flow.
  • Developer ecosystem: Platforms including Agora, Fishjam, LiveKit, Pipecat, Vercel, and Vision Agents integrate the Gemini Live API to manage real-time media streaming infrastructure. Partners such as Salesforce, Genspark, and Lumeris have highlighted the models’ latency, fluidity, and tool-calling capabilities.
  • Safety: All audio generated by these models is watermarked with SynthID, an imperceptible watermark woven directly into the audio output to help prevent misinformation.
  • Pricing context: The underlying Gemini 3.8 Flash model (which powers the Live variants) is priced at $0.75 per million input tokens and $3.75 per million output tokens — the same introductory price as 3.7 Flash.

Unknown

  • Independent benchmark replication: The reported scores — 82.6 on Artificial Analysis’ Speech to Speech Quality Index, 68.6% on τ-Voice, 35.1% on Sierra’s τ-Voice-banking, and 97.7% on Big Bench Audio — come from Google’s own evaluations. No independent third-party verification is cited in the announcement.
  • Real-world latency and cost at scale: While Google emphasizes competitive pricing and low latency, the announcement does not publish SLA-grade latency figures, token consumption profiles for extended reasoning sessions, or per-minute voice pricing for production deployments.
  • Rollout completeness: Several enterprise and Workspace channels are listed as “coming soon” without specific dates. The private preview scope for Gemini Enterprise and Gemini Enterprise for Customer Experience is not quantified.
  • Comparison chart conditions: The Speech to Speech Index chart compares Google’s Live models to GPT-Live-1 Astra (Medium) and Grok Voice Think Fast 2.0 (High). Google does not publish the full prompt set, network conditions, or how those rival builds were configured relative to publicly available endpoints.

Our take

Google is betting the next voice-agent frontier is orchestration — letting a model reason, call tools, and speak simultaneously without the dead air that breaks trust. Extended Thinking’s verbal acknowledgments and live narration directly address that gap. But the vendor-reported benchmarks describe a controlled environment; production workloads with variable networks, accents, and long-running agentic loops will be the real test. Enterprises should treat “coming soon” Workspace and Customer Experience integrations as roadmap signals, not deployed capabilities, and budget for the token overhead that comes with higher thinking levels.

Series: 1. Koray Kavukcuoglu Named DeepMind SVP as Hassabis Becomes Chair · 2. Google DeepMind launches Gemini 3.8 Live and 3.8 Live Extended Thinking for real-time voice agents · DeepMind Leadership

Sources