SOURCE-LINKED INTELLIGENCE
OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation
Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visu
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-15T00:25:28.000Z
First collected: 2026-09-20T09:01:24.920Z. This is not the publication date.