AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Teacher Guidance Misleads: Reward-Aligned On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Aug 28, 2026 · UTC

On-policy distillation (OPD) has recently emerged as a popular post-training paradigm for large language models (LLMs), providing an efficient way to transfer the knowledge and capabilities of teacher models into student models. However, teacher guidance on student-generated prefixes is not always reliable. Training should optimize the model to generate responses that are more likely to be correct, or equivalently, to get higher outcome rewards. But during OPD, the teacher model may provide guidance that discourages the student from moving toward correct trajectories or moves the student towar

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:21:55.975Z. This is not the publication date.