source: Google DeepMind: Intelligent transcription with Gemini 3.5 Transcribe
level: technical
google deepmind released gemini 3.5 transcribe, a speech-to-text model for real-time and recorded audio. it converts raw audio into polished, formatted text, handling background noise, jargon, and disfluencies. developers can access it through the gemini api in google ai studio and gemini enterprise agent platform. two endpoints are available: a live api for sub-second streaming and an interactions api for pre-recorded audio with speaker attribution and word-level timestamps.
the model reports an average word error rate of 4.0% for streaming and 2.6% for non-streaming, as measured by artificial analysis. time to final transcription improves by 70% over the previous chirp 3 model. it supports over 85 languages, custom vocabulary, and multi-speaker identification for up to three speakers. limitations include experimental support for more than three speakers and public preview status, meaning it is not yet generally available for all production uses.
gemini 3.5 transcribe integrates into google products like gboard, the gemini app on macos, and chrome. it enables voice commands that call other gemini models for tasks like image generation or file analysis. developer platforms such as livekit, langchain, and vercel already support the live api. this release follows google's pattern of embedding transcription into everyday tools, aiming to reduce friction for voice-driven workflows in both consumer and enterprise settings.
why it matters: lower word error rates and faster transcription can improve voice agents, captioning, and analytics for ai applications.
source: Google DeepMind: Intelligent transcription with Gemini 3.5 Transcribe