SOURCE-LINKED INTELLIGENCE
AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models
Omni models can describe video content, but can they locate events in time, preserve event order, and judge audio-visual synchronization? We introduce AVTrace (Audio-Visual Temporal Reasoning Assessment and Capability Evaluation), a silver-standard diagnostic suite spanning onset and span grounding, synchronization, next-step prediction, cross-modal localization, chain parsing, and event-conditioned comprehension. It contains 34,114 training examples and category-balanced development and test splits of 3,500 and 7,000 examples. We evaluate five open omni models under their respective input con
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T09:58:27.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T09:58:27.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.