让扑克智能体学会利用对手弱点,超越纳什均衡表现
AlphaExploitem: Going Beyond the Nash Equilibrium in Poker by Learning to Exploit Suboptimal Play

- 用分层Transformer分析历史牌局,动态调整策略
- 对分布内和分布外弱手均有效exploit,且不降低对均衡对手表现
- 适合研究博弈推理与对抗性强化学习的读者
扑克是长期用于测试不确定性下决策能力的不完美信息博弈。为在纳什均衡之上进一步提升收益,智能体可偏离均衡策略以利用对手的次优行为。我们提出AlphaExploitem,通过引入分层Transformer编码器实现对过往牌局的推理,并结合多样化的可 exploited 对手池改进训练过程,从而提升 exploit 能力。在两个标准不完美信息博弈基准上进行训练与评估,实验证明,AlphaExploitem 能有效利用分布内与分布外弱手的失误,同时保持对纳什均衡对手的性能不下降。
原文摘要 · Abstract (English)
Poker is an imperfect information game that has served as a long-standing benchmark for decision-making under uncertainty. To maximize utility beyond the Nash equilibrium, an agent can deviate from Nash-equilibrium policies to exploit suboptimal play. We introduce AlphaExploitem, which extends the competitive RL poker agent AlphaHoldem by using a hierarchical transformer encoder that enables reasoning over previously played hands and modifying the training procedure with the inclusion of a diverse pool of exploitable opponents to facilitate learning to exploit. We train and evaluate AlphaExploitem on two standard benchmarks for imperfect-information games. Empirically, AlphaExploitem successfully exploits weak play by both in- and out-of-distribution opponents, without losing performance against NE opponents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。