Google DeepMind has released Gemini 3.8 Live with Live Avatar, an enterprise feature that adds near real-time video generation to its existing live dialogue models. The capability pairs streaming video with speech, so a virtual agent can listen, see and answer through a dynamic visual persona.

The avatar lip-syncs to what it says, carries natural expressions, and keeps turn-taking fluid rather than stalling between replies. DeepMind says the system processes visual and audio inputs simultaneously to produce what it describes as more natural multimodal conversations. The feature arrives a week after Google's Gemini 3.8 Live launch and builds directly on those dialogue models.

Live Avatar is available from September 24, 2026 in Gemini Enterprise, Google's enterprise tier for business customers. Custom avatars can be generated from a high-quality reference image while preserving likeness, brand styling, or character identity, but that path is gated behind enterprise allowlisting. Asynchronous tool calling lets the agent fetch data in the background while dialogue continues, so a back-end lookup does not force dead air. All output is watermarked with SynthID, DeepMind's imperceptible audio and video watermark.

The feature supports 97 languages, and both lip-sync and expressions adapt when a conversation switches language mid-stream. DeepMind says switching languages does not degrade video fidelity or introduce visual drift. The announcement was posted on behalf of the Gemini Audio Team, with the blog crediting research scientist Shuo-yiin Chang and software engineer CJ Zheng.

Confirmed

  • Product and availability: Gemini 3.8 Live with Live Avatar, available from September 24, 2026 in Gemini Enterprise, Google's enterprise tier
  • Core capability: near real-time video generation paired with live dialogue, giving virtual agents lip-syncing, natural expressions and fluid turn-taking; visual and audio inputs are processed simultaneously
  • Language support: 97 languages with native multilingual speech-to-speech synchronization; lip-sync and expressions adapt mid-conversation without degrading video fidelity
  • Secondary features: custom avatars from high-quality reference images (enterprise allowlisting only), asynchronous tool calling for background data fetching during active dialogue, and SynthID watermarking on all audio and video output
  • Primary source: DeepMind blog post introducing the feature, posted on behalf of the Gemini Audio Team

Unknown

  • Pricing for the Gemini Enterprise tier and any per-seat or usage costs for Live Avatar
  • Whether Live Avatar and its API are offered outside Gemini Enterprise
  • Latency figures for video generation and end-to-end response time
  • Independent benchmark comparisons with competing avatar or video generation services
  • Timeline for opening custom avatar allowlisting beyond initial enterprise access

Our take

Google is racing to productize live multimodal agents for enterprise workflows — customer service, interactive walkthroughs, and brand-specific virtual representatives. The 97-language claim and asynchronous tool calling are strong on paper, but latency in production is the real test, as is whether SynthID watermarking survives adversarial stripping. Pricing is undisclosed, so it is unclear whether Live Avatar carries a premium over the base Gemini Enterprise subscription.

Series: 1. Koray Kavukcuoglu Named DeepMind SVP as Hassabis Becomes Chair · 2. Google DeepMind launches Gemini 3.8 Live with Live Avatar for enterprise · DeepMind Leadership

Sources