AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amp

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:02:11.508Z. This is not the publication date.