AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Online Video Agent Harness for Long Video Understanding

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

Long video understanding often behaves like a visual needle-in-a-haystack problem: query-relevant evidence is sparsely distributed across long temporal spans, while packing dense frames into a single VLM context incurs \textit{context rot} and high cost. Existing video agents often rely on query-agnostic offline preprocessing or ad hoc tool sets, which can miss query-specific details and waste computation. In this work, we present VideoXAgent, a purely online video-agent harness for long video understanding that starts from the given video file and user query, plans and decomposes the task, in

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:22:04.777Z. This is not the publication date.