用神经网络提升博弈算法收敛速度和对抗能力
Deep (Predictive) Discounted Counterfactual Regret Minimization
- 基于价值网络采样优势,通过自举拟合累积优势
- 引入折扣与截断操作,模拟先进博弈算法更新机制
- 在德州扑克等大博弈中收敛更快、对抗更强
对策后悔最小化(CFR)是一类高效求解不完美信息博弈的算法。为提升其在大规模博弈中的适用性,研究者常使用神经网络近似其行为。然而,现有方法多基于基础版CFR,难以有效整合更先进的变体。本文提出一种高效的无模型神经CFR算法,克服了现有方法在逼近高级CFR变体时的局限性。每轮迭代中,该方法基于价值网络收集方差缩减的采样优势,通过自举方式拟合累积优势,并施加折扣与截断操作,以模拟先进CFR变体的更新机制。实验表明,相较于现有的无模型神经算法,该方法在典型不完美信息博弈中收敛更快,在大型德州扑克游戏中展现出更强的对抗性能。
原文摘要 · Abstract (English)
Counterfactual regret minimization (CFR) is a family of algorithms for effectively solving imperfect-information games. To enhance CFR's applicability in large games, researchers use neural networks to approximate its behavior. However, existing methods are mainly based on vanilla CFR and struggle to effectively integrate more advanced CFR variants. In this work, we propose an efficient model-free neural CFR algorithm, overcoming the limitations of existing methods in approximating advanced CFR variants. At each iteration, it collects variance-reduced sampled advantages based on a value network, fits cumulative advantages by bootstrapping, and applies discounting and clipping operations to simulate the update mechanisms of advanced CFR variants. Experimental results show that, compared with model-free neural algorithms, it exhibits faster convergence in typical imperfect-information games and demonstrates stronger adversarial performance in a large poker game.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。