AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

DSE-VTG: Dual-Side Enhancement for Training-Free Video Temporal Grounding

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Text-guided Video Temporal Grounding (VTG) aims to localize the relevant segments in an untrimmed video based on text queries, yet collecting dense temporal annotations and training task-specific models remain costly and brittle under distribution shift. Recent training-free VTG approaches mitigate this issue by directly matching pretrained vision-language representations, but they still face two fundamental information bottlenecks: frame-wise visual encoding overlooks temporal dynamics, while fixed query embeddings cannot resolve query ambiguity. To address these issues, we propose DSE-VTG, a

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:02:11.508Z. This is not the publication date.