AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Trajectory Learnability for Offline On-Policy Distillation with Imperfect Teachers

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

Offline on-policy distillation gains efficiency by collecting student trajectories and teacher supervision once and reusing them throughout optimization. The same reuse makes imperfect supervision persistent. Since even strong teachers can fail, we ask \emph{what remains learnable from imperfect teacher supervision?} Teacher failure is only a coarse problem-level signal and does not imply that all supervision along the associated student trajectory is unhelpful. A natural alternative is to estimate teacher recoverability along the trajectory, but repeated continuations largely erase the effici

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T08:01:03.945Z. This is not the publication date.