AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the context for subsequent text predictions. Given identical speech inputs, we observe markedly lower answer accuracy for the internal text generated in speech-to-text-and-speech (S2TS) mode than for speech-to-text (S2T) responses. We term this discrepancy the \emph{output-mode gap} (OMG). To reduce OMG, we propose \emph{Join

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T11:41:07.830Z. This is not the publication date.