AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models

arXiv · AI, language, vision and robotics · article · Aug 25, 2026 · UTC

Audio clustering is a fundamental task for organizing rapidly growing speech collections, supporting applications such as conversational analysis and speech-driven discovery. However, existing methods rely on fixed acoustic similarity metrics or ASR-based text pipelines, limiting their ability to reorganize the same audio collection under different user-specified perspectives, especially when clustering depends on both linguistic and paralinguistic cues. We introduce audio multi-perspective clustering, where a model directly partitions speech recordings according to a natural-language perspect

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T09:42:05.193Z. This is not the publication date.