AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

CLAMP: Constrained Decoding for Vision-Language Embodied Planning

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executable. A VLM may refer to objects that are not visually observed, select actions whose required affordances are unavailable, or violate syntax and action constraints. We introduce CLAMP, a multimodal constraint-grounding framework that turns scene evidence into decoding-time constraints for a frozen VLM planner. CLAMP uses the initial observation to restrict object references to those supported by th

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:02:11.508Z. This is not the publication date.