source: Google DeepMind: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

level: technical

google deepmind introduced gemini 3.8 live and gemini 3.8 live extended thinking, two new voice models for real-time dialogue. the base model focuses on fast, fluid conversations with visual grounding and automatic language switching across 97 languages. the extended thinking version adds multi-step reasoning and background tool use while speaking. both are available today through the gemini api, google workspace, and the gemini app, with enterprise access in private preview.

gemini 3.8 live extended thinking ranks first on artificial analysis' speech to speech quality index with a score of 82.6. it achieves 68.6% on the tau-voice agentic task benchmark and 35.1% on sierra's tau-voice-banking benchmark. the model scores 97.7% on big bench audio. gemini 3.8 live placed second in the speech agent arena. on servicenow's eva-bench, both models push the pareto frontier for complex voice workflows, balancing accuracy with conversational quality.

the models process visual inputs in near real-time and execute api calls in the background without interrupting speech. extended thinking uses verbal cues like "let me check that" and live progress narration. audio outputs include synthid watermarking for detection. developer platforms such as livekit, langchain, and vercel support the live api. partners like salesforce and genspark highlight low latency and tool calling. the release targets production voice agents for enterprises and consumer apps.

why it matters: these models let developers build voice agents that handle complex tasks while talking naturally, reducing latency and cost for real-time ai applications.


source: Google DeepMind: Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking