arXiv:2608.09586cs.AIcs.GT2026-08

用延续价值优化策略,比传统ICM模型更精准提升扑克锦标赛胜率。

ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs

论文配图:ICM Out! Better Tournament Strategy from Computed Continuations, vs. Solvers and LLMs
图 1 · 摘自论文原文
  • 通过枚举手牌结果并用有限锦标赛模型计算后续状态价值来优化策略
  • 相比固定ICM策略,平均每局多赚214.33美元奖池权益,决策频率偏差达14.08%
  • 在对抗求解器、LLM及阈值玩家时均表现更优,验证方法普适性

独立筹码模型(ICM)将锦标赛筹码转化为奖金权益参考值,策略通常基于此构建。但ICM仅依赖筹码量,忽略行动顺序、盲注义务和座位轮换,也未量化大筹码对短筹码的淘汰压力。本文提出战略延续优化(SCO),通过枚举当前手牌结果,映射至后续状态,并利用有限锦标赛模型计算延续价值,从而优化并冻结当前手牌策略。与固定ICM对比的策略仅改变状态定价方式,其他条件一致。在含100万美元奖金池的三人码/弃牌锦标赛中,分析显示,相较于固定策略,解析式ICM在2,838个状态-座位组合中平均绝对价值误差达9,036美元。该误差显著改变定价范围:相对于自身固定ICM加注范围,SCO使加注频率平均变化14.08%。在946个状态下,保持对手和延续评估器不变,仅替换焦点策略,结果显示SCO平均每局多获214.33美元奖池权益,且在2,433/2,838组匹配单元中占优。该优势在替换求解器对手为两个LLM或一系列非建模阈值玩家后依然成立。该价值-策略-成本链条直接揭示了ICM在锦标赛策略构建中的不足。

原文摘要 · Abstract (English)

The Independent Chip Model (ICM) converts tournament chips into reference prize equity, and policies are routinely constructed against those values. Because ICM reads only stack sizes, it omits action order, blind obligations, and seat rotation, and it does not price the elimination pressure a big stack puts on the short stacks it can bust. Those omissions can alter the successor-state contrasts that determine a move. We introduce Strategic-Continuation Optimization (SCO), a policy-construction method that enumerates current-hand outcomes, maps them to successor states, prices those states with continuation values computed from the finite tournament model, and optimizes and freezes the resulting current-hand policy. The fixed-ICM comparison policy changes one thing only: the same optimizer solves the same game with successor states priced by analytic ICM, so the two policies differ only through that pricing. We evaluate the resulting policies in a three-player jam/fold tournament with a \$1M prize pool. Relative to the frozen strategic-continuation benchmark, analytic ICM has \$9{,}036 mean absolute value error across all 2,838 state--seat entries. That value error rewrites the ranges it prices: measured against each decision point's own fixed-ICM jam range, SCO moves the jam frequency by an average of 14.08\%. To price those different moves, we compare all 946 states and three policy owners while changing only the focal policy and holding both opponents and the continuation evaluator fixed. The policy produced by SCO earns \$214.33 more prize equity per hand on average and is favored in 2,433 of 2,838 matched units. The ordering survives replacing the solver-built opponent with two LLMs and with a family of non-modeling threshold players. This value-to-policy-to-cost chain shows directly when ICM becomes an inadequate objective for tournament strategy construction.

扑克博弈策略优化博弈论强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。