AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Contiguity, Not Importance: Budgeted Repair of Stale KV Caches After Document Edits

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

KV-cache reuse can reduce inference cost in retrieval-augmented generation and agentic systems, but cached contexts may become stale when retrieved knowledge, working memory, or user state is edited. Under causal self-attention, even a local edit can affect downstream KV states. A full re-prefill reliably restores consistency but is costly, whereas refreshing only the edited span can leave downstream dependencies stale. We formulate in-place repair as budgeted recomputation and compare training-free position-selection policies on a factual RAG benchmark with matched direct and derived edits. A

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:20:57.646Z. This is not the publication date.