AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

TempoGround: State-Aware Streaming Visual Grounding with Vision-Language Models

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Visual grounding maps language referents to spatial targets and is central to open-vocabulary perception with vision-language models. Existing methods have made substantial progress on single-frame and video-based visual grounding, yet under streaming inputs they still suffer from identity drift, cross-frame inconsistency, and fragile localization under partial occlusion. To address these issues, we present TempoGround, a VLM-native framework that detects cross-frame object correspondence and explicitly models object presence states, thereby enabling accurate and consistent visual grounding un

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:32:15.665Z. This is not the publication date.