AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

REVA: Reusable Evidence View Aggregation for Context-Efficient RAG Serving

arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC

Retrieval-augmented generation (RAG) improves knowledge-intensive large language model (LLM) applications by conditioning generation on retrieved documents, but longer contexts increase latency, key-value (KV) cache memory, and token cost. Post-retrieval compression can reduce this cost, yet existing compressors often operate independently for each query, rely on auxiliary models or rewriting, and introduce online overhead that can offset the benefit of shorter prompts. We revisit RAG compression from a data-mining perspective by aggregating historical query--document--model interactions into

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:12:12.556Z. This is not the publication date.