SOURCE-LINKED INTELLIGENCE
DIDO: Distilling Interaction-Centric Dynamics into One-Step Denoising for World Action Models
World Action Models (WAMs) use video generation models to predict future visual dynamics for robotic manipulation, but iterative denoising introduces additional latency for closed-loop control. We empirically find that visual content converges at different rates during denoising. Static background structure forms early, whereas the gripper and manipulated object remain blurry after the first step, with their interaction dynamics emerging only through subsequent denoising. Consequently, naively truncating a multi-step video model to one step preserves scene structure but loses the interaction-c
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-14T13:46:57.000Z
First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.