AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

TwinKV: A Composable Repair Pass for KV Cache Eviction via Pairwise Key Redundancy

arXiv · AI, language, vision and robotics · article · Aug 27, 2026 · UTC

Long-context inference is bottlenecked by the memory footprint of the key-value (KV) cache, especially for small models under tight resource budgets. Existing KV cache eviction methods score tokens using the model's attention distribution or, in attention-free variants, each key's distance from a global reference point. Using a controlled leave-one-out probe, we find that attention magnitude is unrelated to a token's causal contribution to the answer (Spearman $ρ=-0.004$), challenging the premise behind dominant eviction methods. We introduce TwinKV, a training-free, attention-free redundancy

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:32:02.028Z. This is not the publication date.