AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Spillover-Aware Multi-Value Steering for Pluralistic LLM Alignment

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

Activation steering controls LLM behavior at inference time by adding learned directions to hidden states, but existing methods handle one concept at a time. Pluralistic alignment, where different stakeholders need different value emphases, requires steering multiple dimensions simultaneously. We show that naive steering produces substantial spillover: the effect intended for one value leaks into others. This parallels the treatment-versus-spillover decomposition in causal inference. We trace spillover to geometric entanglement of steering directions, captured by their Gram matrix, and derive

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:52:07.471Z. This is not the publication date.