SOURCE-LINKED INTELLIGENCE
Deep Noir: Autonomous Steering Discovery via Architectural Chronometry in Transformer Models
Activation steering modifies LLM behavior at inference time, but identifying where and how strongly to steer remains manual. We introduce Deep Noir, a framework that uses Logit Lens convergence and causal head-level attribution to autonomously discover optimal steering parameters. Across three scales (1B x 3, 2-3B x 2, and 7-9B x 4), our engine achieves 16.7 percentage-point improvement on spam at 1B (standard deviation 4.7; 39 runs), with gains increasing to 21 to 42 percentage points at 7-9B across four architectures. On SST-2 sentiment, it achieves a 13.1 percentage-point improvement with z
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-17T17:16:31.000Z
- arXiv · AI, language, vision and robotics · 2026-09-17T17:16:31.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.