AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VLX-VR: An Agentic-Aware Video Reasoning Model

arXiv · AI, language, vision and robotics · article · Sep 9, 2026 · UTC

Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a fixed video context and single-pass inference, limiting adaptive evidence acquisition when observations are incomplete, ambiguous, or conflicting. We present VLX-VR, an agentic-aware video reasoning model trained within a video reasoning framework defined by a Think--Memory--Observation loop. At each step, VLX-VR determines the needed evidence, invokes read_memory or write_memory, incorporates the returned Observation, and decides whether to con

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:32:24.350Z. This is not the publication date.