SOURCE-LINKED INTELLIGENCE
Reach or Solve? Attributing Agentic RL Gains with Checkpoint Handoffs
Reinforcement learning now trains language-model agents that act over dozens of steps in live environments. The gains are large, and they are read as better decision-making. An agent in a closed loop writes its own inputs. Each observation follows from its own earlier actions, so the states it meets late in an episode are partly of its own making. An SFT checkpoint and an RL checkpoint are then scored from different states, even on identical tasks. Endpoint success mixes two changes: where the agent arrives, and what it does once it is there. Restricting the comparison to states both policies
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T03:27:53.000Z
- arXiv · Artificial Intelligence · 2026-09-17T03:27:53.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.