arXiv:2605.14379cs.LGcs.AI2026-05被引 1

通过历史数据初始化游戏起点,加速复杂博弈的探索效率。

Data-Augmented Game Starts for Accelerating Self-Play Exploration in Imperfect Information Games

  • 从人类高手数据中采样中间状态作为训练起点,引导智能体高效探索。
  • 在固定计算预算下,使策略梯度方法在长周期博弈中实现更低可被利用性。
  • 适用于需要高效探索的复杂对抗性游戏研究,如星际争霸、Dota。

大规模不完美信息博弈(如《星际争霸》《Dota》《反恐精英》)寻找近似均衡仍面临计算瓶颈,主要因奖励稀疏且需长期探索。本文提出多智能体起始状态采样策略,显著加速正则化策略梯度方法在双人零和博弈中的在线探索。基于“人类高手的离线示范能覆盖均衡相关高层策略”的假设,我们从离线数据中采样中间状态作为强化学习数据收集的起点,以促进对战略相关子博弈的探索。该方法称为数据增强游戏起点(DAGS)。在合成数据及可分析的长周期控制变体——双人库恩扑克、古夫斯佩尔和一个惩罚隐藏信息偏见信念的反例游戏中进行实验。在固定计算预算下,DAGS使正则化策略梯度方法在探索难度更高的游戏中达到更低的可被利用性。我们发现,在求解不完美信息博弈时,若仅靠数据增广起始状态分布,可能引发偏差均衡,并提出通过多任务观察标志实现简单缓解。最后,我们发布一套新基准环境,大幅提高现有OpenSpiel游戏的探索挑战性和状态数量,同时保持可被利用性测量的可解析性。

原文摘要 · Abstract (English)

Finding approximate equilibria for large-scale imperfect-information competitive games such as StarCraft, Dota, and CounterStrike remains computationally infeasible due to sparse rewards and challenging exploration over long horizons. In this paper, we propose a multi-agent starting-state sampling strategy designed to substantially accelerate online exploration in regularized policy-gradient game methods for two-player zero-sum (2p0s) games. Motivated by an assumption that offline demonstrations from skilled humans can provide good coverage of high-level strategies relevant to equilibrium play, we propose the initialization of reinforcement learning data collection at intermediate states sampled from offline data to facilitate exploration of strategically relevant subgames. Referring to this method as Data-Augmented Game Starts (DAGS), we perform experiments using synthetic datasets and analytically tractable, long-horizon control variants of two-player Kuhn Poker, Goofspiel, and a counterexample game designed to penalize biased beliefs over hidden information. Under fixed computational budgets, DAGS enables regularized policy gradient methods to achieve lower exploitability in games with significantly more challenging exploration. We show that augmenting starting state distributions when solving imperfect information games can lead to biased equilibria, and we provide a straightforward mitigation to this in the form of multi-task observation flags. Finally, we release a new set of benchmark environments that drastically increase exploration challenges and state counts in existing OpenSpiel games while keeping exploitability measurements analytically tractable.

博弈论强化学习自对弈数据增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。