AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

FIVE-VLA: Fast and EffectIVE Autonomous Driving with Recurrent Action Memory

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

State-of-the-art vision-language-action models (VLA) for autonomous driving face critical limitations: excessive parameter counts, inefficient high-resolution image processing, and lack of temporal memory. We introduce Fast and EffectIVE VLA (FIVE-VLA) to address these through two key contributions. First, we employ an efficient vision encoder that processes high-resolution ($448 \times 896$) images while generating only 98 tokens, over $5\times$ fewer than existing approaches, and bypass text generation entirely for single-pass trajectory prediction. Second, we propose Recurrent Action Memory

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:01:03.945Z. This is not the publication date.