SOURCE-LINKED INTELLIGENCE
Efficient Nash Equilibrium Computation for Cybersecurity Games
Computing Nash equilibria of simulation-based cybersecurity games with policy-space response oracles (PSRO) is bottlenecked by payoff estimation: every payoff-matrix entry costs Monte-Carlo rollouts of a slow simulator, while policies and restricted-game solves are cheap. We introduce Regret-Weighted Payoff Sampling (RWPS), a budgeted estimator that simulates only the cells an equilibrium is sensitive to and fills the rest with a surrogate trained on every entry simulated earlier in the run. The sup-norm error bound cannot evaluate such an estimator, because it is set by the cells left deliber
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-16T20:29:07.000Z
- arXiv · Artificial Intelligence · 2026-09-16T20:29:07.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.