SOURCE-LINKED INTELLIGENCE
AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation
Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive polic
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-16T15:10:04.000Z
First collected: 2026-09-19T20:28:26.698Z. This is not the publication date.