SOURCE-LINKED INTELLIGENCE
Causal Localization of the Refusal Direction in Audio Language Models
A large audio language model (LALM) attaches a speech front end to a text language model (LM) that is already safety-aligned. When such a model refuses a harmful spoken request, is the refusal carried by the front end, or inherited from the text LM? We test this with causal interventions. At each model's audio-to-LM interface and at tested LM residual layers, we fit a direction separating harmful from benign prompts, ablate its component, and measure the resulting change in the model's first-token refusal margin. Four of the five models are evaluated under held-out category shift. Across five
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-06T17:29:43.000Z
First collected: 2026-09-25T15:02:49.789Z. This is not the publication date.