AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents

arXiv · AI, language, vision and robotics · article · Sep 17, 2026 · UTC

Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.