source: Google DeepMind: Gemini Robotics 2 brings whole body intelligence to robots

level: technical

Google DeepMind announced Gemini Robotics 2, a set of models that give robots whole-body control, advanced dexterity, and multi-robot collaboration. The system includes a vision-language-action model for motor control, an embodied reasoning model for planning, and an on-device variant for local execution. Robots can now walk, crouch, and manipulate objects to complete tasks like cleaning a cluttered room, marking a shift from pre-programmed motions to learned, adaptable behavior.

The main VLA model controls full humanoids and bi-arm robots, achieving medium to high success on whole-body and gripper tasks, though multi-finger dexterity remains challenging. The embodied reasoning model handles multi-step tasks lasting several minutes and supports teamwork between different robot types. The on-device model adapts to new robot bodies with under 200 examples in a few hours. Safety improvements include a new benchmark for agentic safety and better human proximity detection.

This release builds on earlier Gemini Robotics work, extending control from tabletop tasks to full-body motion. The models are available through Google AI Studio and early-access partnerships. The reasoning model can self-correct and track progress, while the on-device version inherits motion transfer techniques from prior versions. The system aims to move beyond narrow automation toward general-purpose physical AI, though movement speed and dexterous precision still need improvement.

why it matters: It shows how vision-language models can give robots adaptable, whole-body skills and reasoning, reducing the need for task-specific programming and enabling faster deployment across different hardware.


source: Google DeepMind: Gemini Robotics 2 brings whole body intelligence to robots