AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VERPO: Verified Evidence Regularized Policy Optimization

arXiv · AI, language, vision and robotics · article · Sep 5, 2026 · UTC

Verifiable outcome rewards guide language-model post-training, but sequence-level advantages do not identify which token-level decisions should be preserved or revised. Evidence-conditioned Teachers provide denser supervision by replaying sampled trajectories with privileged feedback. Yet indiscriminate imitation risks transferring formatting or reasoning-style shifts that do not support task success. We introduce VERPO, a Verified Evidence Regularized Policy Optimization framework that treats evidence as a proposal for policy correction while retaining the outcome objective. It separates evid

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:32:07.623Z. This is not the publication date.