AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GeoLAM: Learning Geometry-Grounded Latent Actions from Unlabeled Human Videos

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

Human videos provide rich manipulation experience, but extracting action representations that preserve useful motion remains challenging. Visual reconstruction alone can entangle manipulation-related motion with appearance changes and camera movement. We present GeoLAM, a framework for learning geometry-grounded latent actions from action-free human videos. GeoLAM combines future-frame reconstruction through a frozen geometric feature hierarchy with motion supervision from a training-only 4D geometry teacher. The geometric representation provides a structural prior, while the teacher's predict

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:40:59.508Z. This is not the publication date.