SOURCE-LINKED INTELLIGENCE
Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs
Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. We systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length. Using an iso-effect framework that compares capability costs at matched behavi
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-06T08:41:42.000Z
First collected: 2026-09-20T21:12:06.801Z. This is not the publication date.