arXiv:2605.28273cs.AI2026-05中稿 · ICML

用新方法更高效地求解两人零和博弈的均衡,减少迭代次数。

Global Policy-Space Response Oracles for Two-Player Zero-Sum Games

论文配图:Global Policy-Space Response Oracles for Two-Player Zero-Sum Games
图 1 · 摘自论文原文
  • 通过评估扩展后策略集的整体质量来指导策略增长
  • 在多个博弈中实现更低可被利用性,减少约50%迭代次数
  • 适合研究博弈均衡计算与强化学习交叉方向的学者

Policy-Space Response Oracles(PSRO)框架通过深度强化学习迭代扩展受限策略集,将均衡计算推广至大规模零和博弈。核心挑战是在有限算力下构建一个能良好近似全博弈的小型策略种群。现有方法通常基于受限博弈收益计算的元策略进行最优响应扩展,易导致效率低下且全局改进有限。本文提出直接评估扩展后策略集质量来引导种群扩张,采用群体可被利用性(Population Exploitability, PE)衡量受限策略集对全博弈的逼近程度,并设计两阶段探索-选择框架,在扩张过程中显式最小化PE。我们将其实例化为Global PSRO,一种基于DRL的实用算法,通过参数共享的条件神经网络高效生成候选响应并估计PE。在多个双人零和博弈上的实验表明,Global PSRO相较以往方法显著降低可被利用性,以更少的策略迭代逼近纳什均衡。

原文摘要 · Abstract (English)

The Policy-Space Response Oracles (PSRO) framework scales equilibrium computation to large zero-sum games by iteratively expanding a restricted strategy set using deep reinforcement learning (DRL). A central challenge is to construct, under limited computational budgets, a small strategy population whose induced game well approximates the full game. Existing PSRO variants typically expand the population using best responses to meta-strategies computed from restricted-game payoffs, which can lead to inefficient expansions that provide limited global improvement. We propose to guide population expansion by directly evaluating the post-expansion population quality. Specifically, we adopt Population Exploitability (PE) to measure how well a restricted strategy set represents the full game, and introduce a two-phase exploration--selection framework that explicitly minimizes PE during expansion. We instantiate this framework as Global PSRO, a practical DRL-based algorithm that efficiently generates candidate responses and estimates PE via parameter-sharing conditional neural networks. Experiments across multiple two-player zero-sum games show that Global PSRO achieves lower exploitability and approximates Nash equilibria with significantly fewer policy iterations than prior PSRO methods.

博弈论强化学习均衡计算策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。