AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios

arXiv · AI, language, vision and robotics · article · Aug 26, 2026 · UTC

Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not only to interpret video content correctly, but also to satisfy diverse user-specified constraints. Existing benchmarks focus primarily on task accuracy rather than instruction adherence, leaving this capability insufficiently evaluated. To address this gap, we introduce Video-IFBench, a comprehensive benchmark for evaluating instruction following in video understandi

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:22:01.459Z. This is not the publication date.