AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Xiaomi-CocktailASR-1 Technical Report

arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC

Recently, large language model (LLM) based ASR models have achieved significant progress, yet they generally lack support for multi-speaker scenarios, where the cocktail party problem remains a critical bottleneck for further advancing ASR. Existing TS-ASR methods, including end-to-end architectures with speaker embeddings and latest LLM-based explorations suffer from degraded single-speaker performance and the inability to reject when the target speaker is absent. In this paper, we propose Xiaomi-CocktailASR-1, an LLM-based end-to-end TS-ASR architecture. By utilizing reference speech as voic

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:02:05.452Z. This is not the publication date.