Google DeepMind has launched two new text-to-speech models — Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — which it bills as its most expressive audio generation models to date. Developers can use both starting today in Google AI Studio and the Gemini API, with enterprise access via the Gemini Enterprise API coming soon. Consumers will meet Flash TTS inside Gemini Notebook and Flash-Lite TTS inside Google Vids.
Both models move past static voice presets. Gemini 3.8 Flash TTS is aimed at creative direction and character design, letting users build bespoke voices from natural-language prompts (role, accent, characteristics) across more than 100 languages and dialects; it can also replicate a voice from a 30-second audio sample, with consent verification, SynthID watermarking, and C2PA credentials attached. Gemini 3.8 Flash-Lite TTS targets high-volume, cost-efficient work — dubbing, audio content creation, and expressive voice agents — with fine-grained control over tone and pacing.
Confirmed
- Models launched: Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
- Availability (starting September 23, 2026): Developers — Gemini API and Google AI Studio; enterprises — coming soon via the Gemini Enterprise API; consumers — Flash TTS in Gemini Notebook, Flash-Lite TTS in Google Vids.
- Voice design and performance control: Generative voice creation from natural-language prompts across 100+ languages/dialects, a library of 2,000+ production-ready voices including regional varieties, line-by-line direction with stage directions or script cues, long-form generation with minimal speaker drift, native two-speaker scene staging, and scripted vocal bursts and backchanneling (such as laughter, sighs, and gasps, plus interjections like |mhm| and |yeah|).
- Voice replication and safety: 30-second reference sample requiring a verbal consent recording; every generated clip carries SynthID watermarking, plus C2PA credentials.
- Benchmarks (vendor-reported): Flash TTS ranks #1 on the Hume AI Voice Design Benchmark (71.4 overall, 60.8 accent modeling), with Flash TTS and Flash-Lite TTS taking #1 and #2 on the Hume AI Overall Quality Index and top spots in blind human preference evaluations on Voice Arena.
- Partners and roadmap: Agora, LiveKit, Pipecat, and Vercel for developer deployment; Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang for dubbing, localization, and voice agents; voice remixing (prompt-based timbre, pitch, pace, and accent tuning) is listed as coming soon.
Unknown
- Pricing per character or minute for API usage (not disclosed in the announcement).
- Rate limits, latency SLAs, and concurrent request quotas for the Gemini API endpoints.
- Independent replication of the Hume AI and Voice Arena benchmark results by third parties.
- Exact rollout timeline for Gemini Enterprise API access ("coming soon" without a date).
- Whether voice remixing will launch as a separate API or within the existing TTS endpoints.
Our take
Google is treating voice as a programmable design surface rather than a commodity preset. Generative voice creation, line-by-line direction, and built-in consent and watermarking together point at production-grade audio workflows — audiobooks, dubbing, conversational agents — where creative control and IP safety matter equally. The missing pricing and SLA details will decide whether Flash-Lite truly unlocks high-volume use cases or stays a demo-tier offering.
Series: 1. Koray Kavukcuoglu Named DeepMind SVP as Hassabis Becomes Chair · 2. Google DeepMind launches Gemini 3.8 Flash TTS and Flash-Lite TTS with custom voice design and line-by-line dir… · DeepMind Leadership