用微分方程解析Adam-DA在零和博弈中的动态,发现动量作用与优化问题相反。
Understanding Dynamics of Adam in Zero-Sum Games: An ODE Approach

- 通过连续时间微分方程逼近Adam-DA离散动力学
- 发现一阶、二阶动量参数在博弈中作用与优化时相反
- 实验验证该反向动量效应在多模型多数据集上均成立
Adam在训练神经网络中的显著成功,自然推动了其下降-上升变体Adam-DA在零和博弈中的广泛应用。尽管其在实践中广受欢迎,但对Adam-DA的严格理论理解仍滞后。本文推导出描述Adam-DA连续时间极限的常微分方程(ODE),这些方程能精确逼近离散时间下的动态行为,为分析其在零和博弈中的表现提供了可处理的解析框架。利用该方法,我们研究了Adam-DA的两个基本方面:局部收敛性与隐式梯度正则化。分析表明,在零和博弈中,一阶与二阶动量参数的作用恰好与最小化问题中的已知效果相反。通过在多种架构和数据集上的生成对抗网络(GAN)实验,验证了上述预测,展示了该反向动量效应的实际意义。
原文摘要 · Abstract (English)
The remarkable success of the Adam in training neural networks has naturally led to the widespread use of its descent-ascent counterpart, Adam-DA, for solving zero-sum games. Despite its popularity in practice, a rigorous theoretical understanding of Adam-DA still lags behind. In this paper, we derive ordinary differential equations (ODEs) that serve as continuous-time limits of the Adam-DA. These ODEs closely approximate the discrete-time dynamics of Adam-DA, providing a tractable analytical framework for understanding its behavior in zero-sum games. Using this ODE approach, we investigate two fundamental aspects of Adam-DA: local convergence and implicit gradient regularization. Our analysis reveals that the roles of the first- and second-order momentum parameters in zero-sum games are exactly the opposite of their well-documented effects in minimization problems. We validate these predictions through GAN experiments across multiple architectures and datasets, demonstrating the practical implications of this reversed momentum effect.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。