arXiv:2602.10894cs.LGcs.AI2026-02中稿 · ICML被引 1

改进强化学习策略优化,让双人博弈更稳定高效

Revisiting Regularized Policy Optimization for Stable and Efficient Reinforcement Learning in Two-Player Games

  • 结合反KL散度与熵正则化,提升策略更新稳定性
  • 在5个棋类游戏上训练效率超越现有方法
  • 适合研究博弈类强化学习或算法优化的学者

双人零和博弈(如棋类游戏)长期作为强化学习的经典基准。本文重新审视一种结合反Kullback-Leibler正则化与熵正则化的策略优化方法,从理论与实证两方面分析其在双人零和环境中的表现。理论上,研究了策略更新规则在博弈论标准型游戏与有限步长游戏中的稳定性,提供了新的收敛保证,并通过合成游戏的数值实验验证。实证上,基于该方法设计了一种无需模型的强化学习算法,在五种棋类游戏(Animal Shogi、Gardner Chess、Go、Hex、Othello)上进行综合实验,结果表明该智能体在多种环境中均实现更高效的训练。

原文摘要 · Abstract (English)

Two-player games such as board games have long been used as traditional benchmarks for reinforcement learning. This work revisits a policy optimization method with reverse Kullback-Leibler regularization and entropy regularization and analyzes this combination in two-player zero-sum settings from theoretical and empirical perspectives. From a theoretical perspective, we investigate the stability of the policy update rule in two theoretical settings: game-theoretic normal-form games and finite-length games. We provide novel convergence guarantees and verify our theoretical results through numerical experiments on synthetic games. From an empirical perspective, we derive a practical model-free reinforcement learning algorithm based on the regularized policy optimization. We validate the training efficiency of our algorithm through comprehensive experiments on five board games: Animal Shogi, Gardner Chess, Go, Hex, and Othello. Experimental results show that our agent learns more efficiently than existing methods across environments.

强化学习博弈策略优化双人游戏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。