AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation

arXiv · AI, language, vision and robotics · article · Sep 10, 2026 · UTC

Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:02:05.452Z. This is not the publication date.