AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

IMLE-VLA: Fast Single-Step Action Generation for Vision-Language-Action Policies

arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC

Vision-language-action (VLA) policies leverage pretrained vision-language backbones to achieve strong cross-task generalization. A leading design couples this backbone with a dedicated continuous action head trained via diffusion or flow matching. However, such heads rely on iterative multi-step sampling, for example 10 Euler steps in $π_{0.5}$. This creates an inference bottleneck that produces stop-and-go movement in the robot and slower task completion. We introduce IMLE-VLA, which replaces the iterative action head with a single-step conditional generator trained via conditional Implicit M

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:12:12.556Z. This is not the publication date.