AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Inoculation Midtraining with Learned Neologisms

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Large language models (LLMs) often learn both desirable and undesirable properties during post-training. We study whether midtraining, an earlier training stage, can shape which of these properties later generalise. We introduce Inoculation Midtraining, a technique that teaches a base model that unsafe behaviour belongs to a designated context, as indicated by the neologism (a new token) introduced during midtraining, and then post-trains the model on unsafe data within that context. We then evaluate the model outside the context, with the neologism excluded from the system prompt. Across supe

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:41:04.278Z. This is not the publication date.