AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Causal Localization of the Refusal Direction in Audio Language Models

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

A large audio language model (LALM) attaches a speech front end to a text language model (LM) that is already safety-aligned. When such a model refuses a harmful spoken request, is the refusal carried by the front end, or inherited from the text LM? We test this with causal interventions. At each model's audio-to-LM interface and at tested LM residual layers, we fit a direction separating harmful from benign prompts, ablate its component, and measure the resulting change in the model's first-token refusal margin. Four of the five models are evaluated under held-out category shift. Across five

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T15:02:49.789Z. This is not the publication date.