source: Google DeepMind: Introducing Gemini 3.8 Live with Live Avatar
level: business
google deepmind introduced gemini 3.8 live with live avatar, adding near real-time visual presence to its live dialogue models. the feature pairs low-latency streaming video with speech, creating an experience that listens, sees, and speaks with a dynamic visual persona. it includes precise lip-syncing, natural expressions, and fluid turn-taking. available in gemini enterprise, it targets customer service and interactive walkthroughs, letting enterprises expand virtual offerings with more engaging and accessible digital exchanges.
the system processes visual and audio inputs simultaneously and supports asynchronous tool calling, so it can fetch data in the background while continuing dialogue. it handles complex tasks like checking in a hotel guest without interrupting conversation flow. live avatar supports native multilingual speech-to-speech synchronization across 97 languages, dynamically adapting lip-sync and expressions without degrading video fidelity or introducing visual drift. custom avatars can be generated from a reference image, but this requires enterprise allowlisting.
all ai-generated output is watermarked with synthid, an imperceptible watermark woven into audio and video to help detect ai content and reduce misinformation. the feature builds on last week's gemini 3.8 live launch and is available now in gemini enterprise, with api documentation provided. this move reflects a broader push toward multimodal, real-time ai agents for business use, where visual presence and tool use can make interactions feel more natural and capable.
why it matters: real-time avatars with tool use and multilingual support can make enterprise ai agents more engaging and useful for customer-facing tasks.
source: Google DeepMind: Introducing Gemini 3.8 Live with Live Avatar