Speech · body · scene

Puppeteer

Object-Grounded Posture-Aware
Co-Speech Gesture Generation

Natural gestures that listen to speech, respect posture, and understand the physical space around them.

Vida Adeli*1,2,3, Soroush Mehraban*1,2,3, Jacob Rommann1, Harrison Sanborn1, Babak Taati2,3, Cole Clifford1

1Pickford AI 2University of Toronto 3Vector Institute

* equal contribution

Explore

One voice, many physical contexts

Gestures shaped by the world.

Puppeteer outputs across standing and sitting postures, with and without nearby tables, aligned to the same spoken phrases.
The same speech produces semantically aligned motion that adapts to standing, sitting, and nearby furniture. The callouts highlight how scene geometry guides natural hand placement, resting poses, and contact with surrounding surfaces.

Project overview

See Puppeteer in motion.

Abstract

Speech-aligned.
Physically grounded.

Co-speech gestures should do more than follow a voice. They should remain temporally coherent, carry semantic meaning, and fit the body’s physical context.

Puppeteer is a posture-aware, object-grounded gesture diffusion model that operates in a causal latent space. It decomposes long motions into structured primitives, encodes them as temporally ordered latent tokens, and generates gestures conditioned on audio, text, motion history, posture, and surrounding object geometry.

This design supports precise temporal control, gesture completion and in-betweening, while producing diverse, synchronized motion that naturally responds to nearby furniture and other scene constraints.

Method

A three-stage path from speech to scene-aware motion.

Compact causal representations preserve timing; controlled diffusion follows speech and posture; gated fusion grounds the final motion in 3D space.

Three-stage Puppeteer architecture: causal latent gesture primitive space, autoregressive gesture diffusion, and object-aware finetuning.
01

Causal gesture primitives

A causal VAE compresses motion into ordered, continuous latent tokens without discrete quantization.

02

Speech-aware diffusion

Audio windows and text spans give the denoiser fine-grained rhythmic and semantic control.

03

Gated object fusion

Person and object geometry are fused into the model to produce physically grounded gestures.

SceneGes dataset

Gestures that belong in the scene.

SceneGes is a curated synthetic 3D dataset pairing embodied co-speech gestures with surrounding objects. It enables models to learn how passive geometry—tables, chairs, counters, sofas, and more—shapes communicative motion.

26scenarios
153object assets
3Dscene context
SceneGes examples showing people gesturing around varied tables, chairs, counters, a sofa, and a bench.

What Puppeteer enables

Control across time, posture, and space.

Speech synchronization

Windowed audio and word-level text conditioning align motion with rhythm and meaning.

Posture awareness

Generated gestures adapt across standing and varied seated configurations.

Object grounding

3D object geometry informs contact, clearance, gesture amplitude, and resting pose.

Puppeteer showcase

A character, brought to life.

A short visual story of speech becoming expressive, posture-aware, and scene-grounded motion.

Try it

Bring a character to life.

Open the HF demo