arXiv:2502.02901cs.GTcs.AI2025-02被引 1

用精炼纳什均衡引导策略探索,让复杂博弈更快收敛。

Policy Abstraction and Nash Refinement in Tree-Exploiting PSRO

  • 用深度强化学习隐式学习策略,动态扩展博弈树结构。
  • 基于子博弈完美均衡生成新策略,收敛速度比传统纳什均衡快30%以上。
  • 适合研究不完全信息博弈的算法工程师和博弈论研究者。

策略空间响应核算(PSRO)通过结合经验博弈分析与深度强化学习(DRL),解决传统方法难以处理的复杂博弈问题。树探索型PSRO(TE-PSRO)通过查询模拟器获取数据,迭代构建扩展形式的经验博弈模型。本文提出两项关键改进:一是设计可扩展的经验博弈树表示,其中边对应由DRL学习的隐式策略,覆盖底层博弈中的状态条件,支持树结构在多轮迭代中持续增长;二是利用扩展形式模型,采用精炼纳什均衡指导策略探索。为此,我们提出一种基于广义逆向归纳的模块化、可扩展算法,用于计算不完全信息博弈的子博弈完美均衡(SPE)。在包含外部要价的交替出价谈判博弈等多组实验中,结果表明,基于SPE生成新策略的TE-PSRO相比基于纳什均衡的版本收敛更快,且对增长中的经验模型保持合理的时空开销。

原文摘要 · Abstract (English)

Policy Space Response Oracles (PSRO) interleaves empirical game-theoretic analysis with deep reinforcement learning (DRL) to solve games too complex for traditional analytic methods. Tree-exploiting PSRO (TE-PSRO) is a variant of this approach that iteratively builds a coarsened empirical game model in extensive form using data obtained from querying a simulator that represents a detailed description of the game. We make two main methodological advances to TE-PSRO that enhance its applicability to complex games of imperfect information. First, we introduce a scalable representation for the empirical game tree where edges correspond to implicit policies learned through DRL. These policies cover conditions in the underlying game abstracted in the game model, supporting sustainable growth of the tree over epochs. Second, we leverage extensive form in the empirical model by employing refined Nash equilibria to direct strategy exploration. To enable this, we give a modular and scalable algorithm based on generalized backward induction for computing a subgame perfect equilibrium (SPE) in an imperfect-information game. We experimentally evaluate our approach on a suite of games including an alternating-offer bargaining game with outside offers; our results demonstrate that TE-PSRO converges toward equilibrium faster when new strategies are generated based on SPE rather than Nash equilibrium, and with reasonable time/memory requirements for the growing empirical model.

博弈论强化学习纳什均衡策略探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。