AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Zero-shot video highlight detection based on text descriptions and synthetic images

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Detecting video highlights, the most informative or engaging moments in a video, is important for applications such as video summarization and content recommendation. We propose a zero-shot framework that combines CLIP, large language models (LLMs), and diffusion models. Given lightweight video metadata, such as a title or category, an LLM generates textual descriptions of likely highlight events. These descriptions are further converted into synthetic visual prototypes using a diffusion model. Textual and visual representations are matched to video frames using CLIP, enabling frame-level high

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.