SOURCE-LINKED INTELLIGENCE
Safety Beyond the Interface: Detecting Harm via Latent States in Large Language Models
Autonomous systems increasingly rely on Large Language Models (LLMs) yet the safety infrastructure surrounding these models introduces latency and compute overhead. This limits utility in resource-constrained, time-critical deployments. Existing external guardrail models remain blind to the model's internal workings, creating a fundamental assurance gap. We ask: does the model already know when the content is harmful? We extract activations from LLaMA-3.1-8B and train lightweight MLP classifier probes (12.6M parameters) to detect harmful prompts. Evaluated on WildJailbreak, Beavertails, and AE
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · Artificial Intelligence · 2026-09-16T22:25:48.000Z
- arXiv · AI, language, vision and robotics · 2026-09-16T22:25:48.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.