source: Google Research: AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

level: research

google research introduced agenthands, a prototype that gives ai assistants expressive, synchronized hand gestures in extended reality. the system uses a large language model to generate gesture events tied to spoken words, then renders animated hands on an xr headset. this lets agents point, mimic actions, and add visual effects like a red glow for warnings. the goal is to move beyond flat screen overlays and make conversations about physical surroundings more natural and engaging.

a user study with 12 participants compared agenthands to a speech-only baseline on two tasks: orchid care and 3d printer operation. results showed significant gains in spatial grounding, action understanding, and warning noticeability, all with p values below 0.05. participants found it easier to locate objects and follow complex steps. one noted that only when the agent showed a lifting gesture did they realize they needed to lift the orchid to drain water. the system also reduced cognitive load.

agenthands builds on google's prior work in human i/o and sensible agent, and it is published at chi 2026. the prototype uses a taxonomy of hand gestures covering handedness, spatiality, temporal dynamics, interactivity, and visual effects. it registers objects via eye gaze and scene reconstruction, then maps llm output to real-time animations. the research points toward more embodied ai assistants that can operate within a user's physical space, not just analyze it.

why it matters: for ai and data science, agenthands shows how language models can be extended to control physical gestures in real time, improving human-ai collaboration in spatial tasks.


source: Google Research: AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR