SOURCE-LINKED INTELLIGENCE
Steering Equilibrium Selection in Regularized Self-Play via the Reference Policy
Regularized self-play -- the family behind DeepNash's Stratego play -- drives a two-player zero-sum policy to a Nash equilibrium by best-responding to a slowly moving, entropy-regularized reference policy $ρ$. When the game has a polytope of value-equivalent equilibria, the regularizer silently breaks the tie: with a uniform reference it selects the maximum-entropy member, the I-projection of $ρ$ onto the Nash set. Can the reference be used to choose the equilibrium on purpose? On five exactly solvable games plus a 2-D polytope, with exact best responses and equivalence tests over independent
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T07:30:23.000Z
- arXiv · Artificial Intelligence · 2026-09-17T07:30:23.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.