探索扑克类博弈中策略的可学习表示,验证了嵌入效果。
Towards Learning Representations of Policies in Two-Player Zero-Sum Imperfect-Information Games

- 构建博弈策略数据集并设计自监督学习方法
- 在克努扑克和莱德克扑克上验证嵌入有效性
- 首个系统比较博弈策略表征学习的方法
我们研究了在双人零和不完美信息博弈中学习有用策略表示(嵌入)的问题。本文提出三项贡献:首先,构建特定游戏的策略数据集生成方法;其次,提出策略表示学习方法;第三,设计下游任务以评估表示效果。我们在克努扑克(Kuhn Poker)和莱德克扑克(Leduc Poker)上评估了每种数据集生成方法、嵌入方法及下游任务。尽管方法基础,但结果表明所学嵌入中存在有效的行为表示。据我们所知,本工作是首个系统比较自监督学习技术在博弈策略表示中的应用。代码已开源,供扩展使用。
原文摘要 · Abstract (English)
We investigate the problem of learning useful policy representations (embeddings) in two-player zero-sum imperfect-information games. We make three contributions: First, we introduce methods of creating datasets of policies for a given game. Second, we propose methods to learn policy representations. Third, we introduce downstream tasks to evaluate the effectiveness of such representations. We evaluate each dataset method, embedding method, and downstream task on Kuhn and Leduc Poker. Although our methods are very basic, we demonstrate that useful behavioral representations are present in the learned embeddings. To our knowledge, this work is among the first to systematically compare self-supervised learning techniques for learning policy representations in games. Our code is available at https://github.com/VitamintK/ssl-project for others to extend.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。