AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Know When to Stop, Where to Restart: Accelerating Multi-Turn Agentic On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

On-policy distillation (OPD) has become a standard approach for transferring capabilities from large teachers to compact students. Its cost, however, is dominated by autoregressive student rollouts and scales poorly in multi-turn agentic settings. Existing acceleration methods truncate or relocate the supervision signal according to fixed, offline budgets, despite substantial variation in teacher-signal reliability both within and across trajectories. Our empirical analysis on $τ^2$-bench reveals a clear structure in this variation: informative supervision is concentrated in the prefix of each

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:21:05.240Z. This is not the publication date.