source: Google DeepMind: Introducing agentic video understanding with Gemini

level: technical

google deepmind has introduced agentic video understanding for gemini 3.7 flash, 3.6 flash, and 3.5 flash-lite. the feature lets models dynamically scan video segments instead of processing at a fixed frame rate. developers can enable it by setting processing to "agentic" in the api configuration. it is available today in google ai studio and the gemini enterprise agent platform for video uploads and youtube videos.

benchmarks show token consumption drops by up to 88% and costs fall by up to 66%, while accuracy improves by up to 7%. gains are largest on long-form video, from 10-minute guides to multi-hour recordings. gemini 3.7 flash with agentic understanding sits at the accuracy-to-cost pareto frontier. the feature uses standard api token pricing with no extra fee, and it will roll out to the gemini app and youtube's ask feature.

the agentic loop lets gemini choose what to watch, at what speed, and through which modality, fetching only needed moments. this supports sub-second moment retrieval, long-form search, anomaly detection, and precise counting. early access partners reported strong performance. the change reduces manual overhead for developers who previously had to implement such logic themselves, making video analysis more practical for large-scale applications.

why it matters: lower token use and cost make long-form video analysis feasible for more ai and data science projects without sacrificing accuracy.


source: Google DeepMind: Introducing agentic video understanding with Gemini