AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

In reinforcement learning for large language models, Proximal Policy Optimization (PPO) commonly uses a critic to estimate state values and reduce the variance of policy updates. However, we uncover a systematic failure mode in PPO critics, which we call Value Flattening: state values, estimated from multiple Monte Carlo continuations, change sharply across intermediate states while critic predictions remain comparatively flat. We further observe this phenomenon in a controlled FrozenLake environment and find that it becomes more pronounced as the state space grows. Our theoretical and empiric

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-19T20:28:26.698Z. This is not the publication date.