AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. We systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length. Using an iso-effect framework that compares capability costs at matched behavi

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:12:06.801Z. This is not the publication date.