SOURCE-LINKED INTELLIGENCE
Stress-testing Alignment Midtraining
When aligning frontier models through post-training techniques, it is not possible to directly demonstrate all of the behaviours we want a model to exhibit in all possible deployment environments; our model must generalise outside of the post-training distribution. One proposed solution is alignment midtraining (AMT), which continues pretraining on large volumes of alignment-relevant documents to encourage generalisation in later stages of training. Despite the prominence of AMT as an alignment approach, there is limited public evidence for its effectiveness. To resolve this, we identify sever
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T13:56:22.000Z
- arXiv · Artificial Intelligence · 2026-09-17T13:56:22.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.