AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context

arXiv · AI, language, vision and robotics · article · Aug 26, 2026 · UTC

Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:22:01.459Z. This is not the publication date.