AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Towards High-DoF Dexterous Manipulation through VLA Post-Training

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

Imitation-learned vision--language--action (VLA) foundation models acquire broad manipulation capabilities by scaling robot data across tasks and embodiments, but reliable deployment on a specific downstream task and hardware platform still requires post-training. Dexterous hands make this adaptation particularly difficult: their broad behavioural repertoire and high degree of freedom create a large and structured action space. Three obstacles are central: open-source VLAs do not natively provide an action interface for high-DoF hands; gesture mismatch during human-gated DAgger takeover create

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:28:21.856Z. This is not the publication date.