arXiv:2608.14761cs.GTcs.LG2026-08

提出确定性框架,分析持久性策略更新在扑克中的收敛性与偏差问题。

CFR without Unbiasedness: Deterministic Guarantees for Persistent Public-Chance Schedules

  • 用确定性转移定理分析持久部分评估的反馈偏差机制。
  • 实测显示固定顺序比重洗更优,32~64次预算后完全覆盖占优。
  • 适用于设计和验证持久性博弈学习策略,尤其对扑克类游戏有指导意义。

在有限公共机会切分下,反事实遗憾最小化(CFR)需决定每次遗憾更新前评估多少结果。精确评估在单一策略配置下处理完整切分;持久部分评估则以无放回顺序遍历不断演化的配置,每轮仅覆盖每个结果一次。尽管每轮覆盖率相同,其反馈通常存在条件偏差,因早期批次影响后续批次所见配置。本文建立针对均匀、非嵌套加法型公共切分的确定性目标转移定理。该定理将完整切分的可利用性上限由交付反馈的遗憾值与一个耦合前缀覆盖差异与实际策略路径运动的公共负债项共同约束。因此,连续平衡调度在预设平均权重下对加法符号遗憾匹配(RM)与RM+均保证收敛;而固定RM+构造证明该差异-路径乘积在一般情况下为必要。定理的分量解析形式可将执行轨迹转化为数值可利用性证书。在两个公开发布的一对一无限注德州扑克转牌末局中,持久顺序显著优于随机重洗,即使每轮覆盖率相同;部分覆盖在所有注册的浅层预算对比中胜出。深度研究发现,在32至64次完整切分结果预算之间存在拐点,之后完全覆盖占据优势。这些结果将公共机会宽度与顺序视为学习变量,并为设计和审计持久性CFR调度提供了确定性基础。

原文摘要 · Abstract (English)

At a finite public-chance cut, counterfactual regret minimization (CFR) must choose how many outcomes to evaluate before each regret update. Exact evaluation processes the full cut at one strategy profile; persistent partial evaluation processes a fixed without-replacement order across evolving profiles. The latter covers every outcome once per epoch, yet its feedback is generally conditionally biased because earlier batches influence the profiles seen by later batches. We establish a deterministic target-transfer theorem for uniform, nonnested additive public cuts. The theorem bounds full-cut exploitability by regret on the delivered feedback and a public-debit term that couples prefix coverage discrepancy with motion along the realized strategy path. Consecutively balanced schedules consequently converge for additive signed regret matching (RM) and RM+ under predetermined averaging weights, while a fixed RM+ construction proves that the discrepancy--path product is necessary in general. A component-resolved form of the theorem converts an execution trace into a numerical exploitability certificate. On two released heads-up no-limit hold'em turn endgames, persistent order improves substantially over fresh reshuffling despite identical epochwise coverage, and partial coverage wins every registered shallow matched-budget comparison. A depth study locates a crossover between 32 and 64 full-cut outcome budgets, after which complete coverage dominates. These results characterize public-chance width and order as learning variables and provide a deterministic basis for designing and auditing persistent CFR schedules.

博弈学习扑克AICFR收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。