AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

As LLM conversations grow to hundreds of turns, full-context injection incurs $O(N^2)$ cumulative token costs, while lossy summarization or hard truncation irreversibly discards historical state. We propose Pull, a session router that maintains an addressable metadata directory via a local, deterministic Purifier (zero LLM calls, millisecond-level latency). At query time, the LLM lazily materializes only the turns it needs; unmaterialized turns remain accessible but collapsed. Unlike irreversible compression, Pull's materialization is reversible; subsequent queries can expand any collapsed tur

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.