负动量提升约束极小极大博弈学习速度,效果远超现有方法。
Rapid Learning in Constrained Minimax Games with Negative Momentum
- 提出负动量缓冲更新框架,扩展至约束极小极大博弈
- 理论证明带熵正则化的算法可收敛,实验性能显著超越基线
- 适用于正常形式与广义形式博弈,适合强化学习和博弈建模场景
本文研究负动量技术在约束极小极大博弈中的应用。从直观的力学视角出发,提出一种新的动量缓冲更新框架,将负动量的成果从无约束设置推广至约束设置,为经典博弈求解算法提供通用增强。此外,我们为带有熵正则化的动量增强算法提供了收敛性理论保证,并将其拓展至广义形式博弈。在正常形式博弈(NFGs)与广义形式博弈(EFGs)上的实验结果表明,所提动量技术能显著提升算法性能,大幅超越原始版本及当前最优基线。
原文摘要 · Abstract (English)
In this paper, we delve into the utilization of the negative momentum technique in constrained minimax games. From an intuitive mechanical standpoint, we introduce a novel framework for momentum buffer updating, which extends the findings of negative momentum from the unconstrained setting to the constrained setting and provides a universal enhancement to the classic game-solver algorithms. Additionally, we provide theoretical guarantee of convergence for our momentum-augmented algorithms with entropy regularizer. We then extend these algorithms to their extensive-form counterparts. Experimental results on both Normal Form Games (NFGs) and Extensive Form Games (EFGs) demonstrate that our momentum techniques can significantly improve algorithm performance, surpassing both their original versions and the SOTA baselines by a large margin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。