arXiv:2511.12083cs.AI2025-11AAAI被引 1

用嵌入空间提升扑克牌博弈策略求解精度与速度

No-Regret Strategy Solving in Imperfect-Information Games via Pre-Trained Embedding

  • 将信息集映射到低维连续嵌入空间,保留细微差异
  • 相同资源下收敛速度更快, exploitability 显著降低
  • 首个在扑克AI中预训练嵌入进行策略求解的算法

大规模不完美信息扩展形式博弈(如无限注德州扑克)中的高质量信息集抽象仍是核心挑战,受限于有限空间资源,难以求解完整游戏策略。现有先进AI方法依赖预训练离散聚类进行抽象,但硬分类会不可逆地丢弃关键信息,特别是信息集间的可量化细微差异,影响策略质量。受自然语言处理中词嵌入启发,本文提出 Embedding CFR 算法,将信息集特征预训练并嵌入至互联的低维连续空间,使向量更精确捕捉信息集间的差异与关联。该算法在嵌入空间中基于悔悟累积与策略更新求解策略,并提供理论分析证明其能减少累计悔悟。在扑克实验中,相同空间开销下,Embedding CFR 的可被利用性收敛速度显著优于基于聚类的抽象算法,验证了其有效性。据我们所知,这是首个在扑克AI中通过低维嵌入预训练信息集抽象以求解策略的算法。

原文摘要 · Abstract (English)

High-quality information set abstraction remains a core challenge in solving large-scale imperfect-information extensive-form games (IIEFGs)--such as no-limit Texas Hold'em--where the finite nature of spatial resources hinders solving strategies for the full game. State-of-the-art AI methods rely on pre-trained discrete clustering for abstraction, yet their hard classification irreversibly discards critical information: specifically, the quantifiable subtle differences between information sets--vital for strategy solving--thus compromising the quality of such solving. Inspired by the word embedding paradigm in natural language processing, this paper proposes the Embedding CFR algorithm, a novel approach for solving strategies in IIEFGs within an embedding space. The algorithm pre-trains and embeds the features of individual information sets into an interconnected low-dimensional continuous space, where the resulting vectors more precisely capture both the distinctions and connections between information sets. Embedding CFR introduces a strategy-solving process driven by regret accumulation and strategy updates in this embedding space, with supporting theoretical analysis verifying its ability to reduce cumulative regret. Experiments on poker show that with the same spatial overhead, Embedding CFR achieves significantly faster exploitability convergence compared to cluster-based abstraction algorithms, confirming its effectiveness. Furthermore, to our knowledge, it is the first algorithm in poker AI that pre-trains information set abstractions via low-dimensional embedding for strategy solving.

博弈论嵌入空间扑克AI策略求解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。