AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Recovering Aggressively Pruned Vision-Language-Action Models with Offline Hidden-State Distillation

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

Vision-language-action (VLA) models let robots follow language instructions, but their language backbones of several billion parameters are the main obstacle to running them on robot hardware. Structured pruning reduces that backbone, and removing 63% of it from OpenVLA-OFT drops LIBERO-Long success from 93.2% to 0.8%. A recent approach restores such a model with supervised fine-tuning followed by reinforcement learning, which needs online rollouts and hundreds of GPU-hours. We recover most of the lost success entirely offline. Width pruning narrows the blocks but keeps the residual stream at

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:28:21.856Z. This is not the publication date.