SOURCE-LINKED INTELLIGENCE
Local Sparsity Enables Unsupervised LLM Safety Detection
Deployment-time safety methods for large language models (LLMs) are predominantly supervised and assume access to unsafe training data. Nevertheless, new attacks and harm categories regularly arise, not captured by models trained in such a supervised fashion. An alternative approach is to view this problem through the lens of anomaly detection, namely, to rely solely on modeling safe data and flagging out-of-distribution inputs. However, LLM activations lie in a high-dimensional space, raising concerns about whether anomaly detection is statistically feasible. We show that, under the linear re
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T12:23:04.000Z
- arXiv · Artificial Intelligence · 2026-09-17T12:23:04.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.