AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SAGE: State-Grounded, Abstention-Aware Evaluation of Task-Oriented Dialogue Agents

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

Evaluating task-oriented dialogue agents requires judging not merely whether a reply reads well but whether each turn advances the underlying workflow state correctly--a distinction conventional holistic LLM judges can miss because they evaluate the available context as a single unit and require one or more full-model calls per turn. We propose SAGE (State-Grounded Abstention-Aware Evaluation), which compiles a workflow specification and per-turn state diff into atomic, schema-grounded criteria and routes each through a cascade of symbolic and encoder/NLI verifiers that abstain rather than gue

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:21:59.299Z. This is not the publication date.