Google DeepMind has launched agentic video understanding across its latest Gemini models — Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite. The feature is available today for uploaded video and YouTube URLs via the Gemini API in Google AI Studio and the Gemini Enterprise Agent Platform.
Static processing samples video at a default 1 FPS. The new agentic mode pairs the model's reasoning with native video tools to search, scan, and inspect target segments across visual frames, audio, and transcripts. Google says this reduces token consumption by up to 88%, cuts analysis costs by up to 66%, and improves accuracy by up to 7% on standard video understanding benchmarks.
What's new
- Models supported: Gemini 3.7 Flash, 3.6 Flash, 3.5 Flash-Lite
- Token reduction: up to 88% (Gemini 3.7 Flash with agentic mode)
- Cost reduction: up to 66% on video analysis
- Accuracy gain: up to 7% on benchmarks such as LongVideoBench
- Availability: starting September 1, 2026, via Gemini API in Google AI Studio and Gemini Enterprise Agent Platform — uploaded video and YouTube URLs
- Pricing: standard Gemini API token pricing, no additional feature fee
- Activation: set
processing: "agentic"in the API configuration
How it works
Instead of a single-pass ingestion at a fixed frame rate, agentic video understanding runs an internal agentic loop. The model decides what to watch, at what speed, and through which modality — frames, audio, or transcript — fetching only the moments and signals needed for the query. Developers previously had to build this logic manually; the agentic loop now handles it internally, significantly reducing development overhead.
Key capabilities
- Sub-second moment retrieval: pinpoints split-second state changes and tight cut boundaries missed at 1 FPS, enabling precise automated video editing.
- Long-form needle-in-a-haystack search: answers complex queries across multi-hour videos without consuming millions of tokens.
- Anomaly detection: resamples interesting time windows at higher FPS to inspect rapid motion and subtle visual artifacts.
- Counting actions and objects: accurately tracks repeated physical movements and distinct objects over time.
Why it matters
The efficiency gains are most pronounced on long-form content — from 10-minute how-to guides to 90-minute lectures and multi-hour recordings — where static processing forces developers to choose between high token costs or techniques that drop critical details. On Minerva, 1H-VideoQA, and LVBench, Google's published chart shows Gemini 3.7 Flash with agentic mode using 33.6K–36.0K tokens per query versus 80.9K–397.6K with static processing (up to 88% savings on long-video sets), while accuracy rises to 79.0%–88.6% versus 73.7%–87.5% — placing agentic 3.7 Flash on the accuracy-to-cost pareto frontier (high thinking, low media res, 1 FPS static baseline).
Google also confirmed the capability will roll out to the Gemini app for all users across Flash and Flash-Lite models soon, and will power YouTube's "Ask YouTube" feature on the video watch page in the coming months, leveraging Gemini to deliver higher-quality answers grounded in the visuals.
Our take
Agentic video understanding moves the bottleneck from token budget to model reasoning quality. By letting the model decide where to look, Google shifts the economics of long-form video analysis on the same Gemini 3.7 Flash SKU it launched in August: developers no longer need to pre-sample or chunk footage, and the same API call can handle a 30-second clip or a three-hour lecture with proportional cost.
Series: 1. Koray Kavukcuoglu Named DeepMind SVP as Hassabis Becomes Chair · 2. Google DeepMind launches agentic video understanding for Gemini 3.7 Flash, 3.6 Flash, and 3.5 Flash-Lite · DeepMind Leadership