SOURCE-LINKED INTELLIGENCE
Can Data Attribution Filter Out Subliminal Learning? Not Reliably
Subliminal learning allows language models to transmit behavioral traits through training data with no obvious semantic relationship to those traits, undermining content-based data filtering as a safety intervention. Training data attribution offers an alternative: it identifies the training examples responsible for a given model behavior, independent of their semantic content, and so may apply in exactly the cases where semantic inspection fails. We evaluate three gradient-based attribution methods (GradCos, a contrastive GradCos variant, and EK-FAC) across three models, comparing them agains
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T10:34:18.000Z
- arXiv · Artificial Intelligence · 2026-09-17T10:34:18.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.