AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

arXiv · AI, language, vision and robotics · article · Aug 26, 2026 · UTC

Co-speech gesture generation has made significant progress toward realistic full-body motion from speaker audio, yet existing models lack fine-grained spatial controllability of individual joints. To address this, we introduce \emph{InteractGesture}, a model-agnostic, inference-time method for spatially controllable gesture generation. \emph{InteractGesture} guides target latent estimates of a diffusion sampler through a differentiable RVQ-VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. A primary challenge in streaming co-speech generation is ch

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:22:01.459Z. This is not the publication date.