AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Reproducing and Evaluating the Generalizability of Subliminal Learning in Open-Weight Models

arXiv · AI, language, vision and robotics · article · Sep 11, 2026 · UTC

In this reproduction paper we investigate subliminal learning, a consequence of distillation where teacher models transmit behavioral preference traits through semantically unrelated data. The original paper explores two types of traits (animal preferences and misalignment), three data modalities (number sequences, code, and chain of thought), and several model families. We reproduce their experiments and extend the setup along three axes: new preference categories (actors and politicians), a new task (chess move generation), and an additional open-weight model (Ministral8B). We also run a con

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T18:22:04.777Z. This is not the publication date.