AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Long autoregressive video generation faces a fundamental memory challenge: with a finite attention window, a model must decide which information from an ever-expanding history to retain. Existing methods organize memory temporally, preserving recent frames while compressing or discarding older ones. We instead propose RECAP-Forcing, organizing memory by appearance novelty. A long video is not merely a sequence of frames, but an evolving cast of subjects, objects, and scenes whose identities must remain consistent over time. We organize memory by retaining the KV cache associated with newly app

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:51:59.673Z. This is not the publication date.