AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Apple Machine Learning Research · article · Sep 11, 2026 · UTC

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth references. This paradigm suffers from the “one-to-many” nature of video description, where high-quality captions are often penalized for lexical mismatches or valid shifts in visual focus. Furthermore, such assessments are typically one-dimensional, failing to provide a fine-grained analysis of caption quality. To address this, we redefine caption quality via information fidelity: A caption must maximize the coverage

Read original source ↗ Open in workspace

recordType
article
region
Global

Evidence & attribution

First collected: 2026-09-19T20:26:46.936Z. This is not the publication date.