AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

How Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE

arXiv · AI, language, vision and robotics · article · Sep 9, 2026 · UTC

Directional ablation removes an aligned language model's ability to refuse by projecting a single "refusal direction" out of the weights that write the residual stream. It needs no gradient-based training and no optimization, only a few hundred contrastive prompts, which makes it the canonical white-box attack on open-weight alignment. However, it has been established only on dense models up to roughly 70B parameters. We study whether it survives the shift to frontier mixture-of-experts (MoE) models whose residual streams are no longer a single tensor and whose weights ship quantized. We apply

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T19:52:05.078Z. This is not the publication date.