AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

GroundingVLN: Reasoning and Acting with Grounding for Vision-Language Navigation

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

Although vision-language models (VLMs) possess strong visual understanding and reasoning capabilities, existing vision-and-language navigation (VLN) agents struggle to connect semantic reasoning with spatial execution. Two coupled gaps remain in this connection, as intermediate reasoning is not explicitly anchored to visual evidence and high-level decisions lack precise spatial goals to guide low-level motion. Cognitive science suggests that human navigation bridges these levels hierarchically by anchoring cognition to relevant landmarks and guiding locomotion toward spatial goals. Motivated b

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:01:03.945Z. This is not the publication date.