AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Concord: A Video Relational Algebra for Cross-Modal Query Optimization

arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC

Semantic video queries let users embed natural language prompts and use multimodal large language models (MLLMs) to interpret the video. Such queries are increasingly popular for querying video data. However, their expressiveness comes at a steep cost: an MLLM may process hours of media to return only seconds of relevant output, making naive execution slow, expensive, and inaccurate. We propose Concord, a system for expressing and optimizing semantic video queries. We makes three contributions. First, we introduce Video Relational Algebra (VRA), a nested algebra over videos, transcripts, frame

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:52:07.471Z. This is not the publication date.