AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this work, we propose an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head. Our approach requires minimal additional parameters and no significant architectural changes to the base

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.