AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

AURA: Unified Multimodal Framework for Conversational Music Editing

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaf

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:41:04.663Z. This is not the publication date.