arXiv:2410.01575cs.GTcs.AI2024-10被引 1

提出H-PSRO框架,解决异构团队零和博弈的全局均衡求解难题。

Computing Ex Ante Equilibrium in Heterogeneous Zero-Sum Team Games

  • 为异构队友设计分层策略参数化,支持顺序优化提升团队收益
  • 在异构矩阵博弈中实现收敛,较传统方法降低显著可被利用性
  • 适用于异构与同质团队场景,是首个适配异构团队的PSRO框架

在异构团队零和博弈中,团队内成员角色不同,现有基于策略空间响应归约(PSRO)的方法如Team PSRO因无法覆盖完整的团队策略空间,导致陷入次优的预期均衡,可被利用性高且不收敛。本文首先为异构队友设计参数化策略,并证明顺序优化能单调提升团队收益。进而提出异构-PSRO(H-PSRO),将顺序相关机制融入PSRO框架,成为首个专为异构团队设计的PSRO方法。理论证明其在异构博弈中可实现更低可被利用性。实验显示,H-PSRO能在传统非异构基线无法求解的异构矩阵博弈中实现收敛;同时在异构与同质设置下均优于现有方法。

原文摘要 · Abstract (English)

The ex ante equilibrium for two-team zero-sum games, where agents within each team collaborate to compete against the opposing team, is known to be the best a team can do for coordination. Many existing works on ex ante equilibrium solutions are aiming to extend the scope of ex ante equilibrium solving to large-scale team games based on Policy Space Response Oracle (PSRO). However, the joint team policy space constructed by the most prominent method, Team PSRO, cannot cover the entire team policy space in heterogeneous team games where teammates play distinct roles. Such insufficient policy expressiveness causes Team PSRO to be trapped into a sub-optimal ex ante equilibrium with significantly higher exploitability and never converges to the global ex ante equilibrium. To find the global ex ante equilibrium without introducing additional computational complexity, we first parameterize heterogeneous policies for teammates, and we prove that optimizing the heterogeneous teammates' policies sequentially can guarantee a monotonic improvement in team rewards. We further propose Heterogeneous-PSRO (H-PSRO), a novel framework for heterogeneous team games, which integrates the sequential correlation mechanism into the PSRO framework and serves as the first PSRO framework for heterogeneous team games. We prove that H-PSRO achieves lower exploitability than Team PSRO in heterogeneous team games. Empirically, H-PSRO achieves convergence in matrix heterogeneous games that are unsolvable by non-heterogeneous baselines. Further experiments reveal that H-PSRO outperforms non-heterogeneous baselines in both heterogeneous team games and homogeneous settings.

博弈论团队协作强化学习PSRO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。