AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Sep 15, 2026 · UTC

Privileged on-policy distillation improves multimodal reasoning by allowing a teacher to evaluate student trajectories using rich, training-only visual evidence. Both models score these trajectories while conditioning on the same student-generated prefix. When a student misinterprets an image early in a response, this accumulating erroneous rationale eventually pulls the teacher away from its visual evidence. The teacher and student converge on the same hallucination, causing standard cross-model supervision to collapse precisely where correction is most needed. We find that the teacher's visu

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T09:01:24.920Z. This is not the publication date.