AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

What Did the MLLM Hear? Token-Level Spectro-Temporal Grounding for Audio MLLM Explainability

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

Audio-based Multimodal Large Language Models (MLLMs) can generate detailed natural-language descriptions of complex acoustic scenes, yet it remains unclear which parts of the input audio support each generated token. This is particularly challenging because acoustic evidence is distributed across time and frequency, and concurrent sound events may overlap temporally while occupying different spectral regions. We introduce STAG, to our knowledge the first post-hoc framework for token-level spectro-temporal grounding of captions generated by audio-based MLLMs. STAG estimates the temporal support

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:22:04.777Z. This is not the publication date.