AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

MarKey: Marginal Utility Guided Greedy Keyframe Selection for Long Video Understanding

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Long-video understanding remains challenging for multimodal large language models (MLLMs) because densely encoding long frame sequences is computationally expensive, while uniform sampling under a limited visual budget can miss sparse yet decisive evidence. Recent training-free keyframe selection methods have enabled more efficient inference and yielded promising performance gains. However, many existing methods score frames largely in isolation without explicitly considering how each candidate complements the currently selected subset, potentially resulting in redundant selections and incompl

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.