SOURCE-LINKED INTELLIGENCE
Improving Cross-embodiment Transfer in Latent Action Models with Action-Similarity Supervision
As generalist robot policies gain vision and language from web-scale pretraining, demonstrations remain costly to collect and tied to the robot that recorded them. Latent action models (LAMs) address both by learning latent actions from action-free videos that can be shared across embodiments, however, in practice, LAMs are sensitive to background visual noise, and the same motion from two different robots may be encoded with different latents. One solution to the background visual noise is to add an auxiliary loss predicting the robot action from the latent action, further associating the lat
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T07:55:35.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T07:55:35.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.