SOURCE-LINKED INTELLIGENCE
Mitigating Retaliatory Algorithmic Collusion in Repeated Games
Reinforcement learning agents trained to maximize their own reward in repeated interactions can converge to supra-competitive outcomes resembling explicit collusion, without communication or shared design. Existing mitigation approaches are largely tied to specific economic settings, like two-sided platforms and auctions, leaving open how to design interventions for general repeated games. We address this gap by formalizing the connection between empirical observations from prior work on Q-learning collusion and classical theory of Simple Penal Codes (SPCs). We show any non-trivial SPC induces
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T15:12:59.000Z
- arXiv · Artificial Intelligence · 2026-09-17T15:12:59.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.