提出新型均值场梯度下降法求解多玩家零和博弈的混合纳什均衡。
Convergence of Time-Averaged Mean Field Gradient Descent Dynamics for Continuous Multi-Player Zero-Sum Games
- 采用带动量和指数时间平均的耦合均值场梯度流
- 在固定熵正则下以指数速度收敛到混合纳什均衡
- 支持不同温度下的模拟退火,适用于无正则化问题
针对具有均值场交互的多玩家零和博弈中混合纳什均衡(MNE)的逼近,本文提出一种针对 $K$ 个玩家($K\geq2$)的均值场梯度下降动力学。玩家策略分布的演化遵循带有动量的耦合均值场梯度流,并引入指数加权的时间平均梯度。在固定熵正则化条件下,证明了该动力学以指数速率收敛至混合纳什均衡,相对于总变差度量。相比先前具有不同平均因子的类似时间平均动力学的多项式收敛率,本方法有显著提升。此外,不同于以往分时尺度的两阶段方法,本文方法将所有玩家类型置于同一时间尺度。通过选择合适的递减温度参数,还证明了该均值场动力学的模拟退火版本可收敛至初始未正则化问题的混合纳什均衡。
原文摘要 · Abstract (English)
The approximation of mixed Nash equilibria (MNE) for zero-sum games with mean-field interacting players has recently raised much interest in machine learning. In this paper we propose a mean-field gradient descent dynamics for finding the MNE of zero-sum games involving $K$ players with $K\geq 2$. The evolution of the players' strategy distributions follows coupled mean-field gradient descent flows with momentum, incorporating an exponentially discounted time-averaging of gradients. First, in the case of a fixed entropic regularization, we prove an exponential convergence rate for the mean-field dynamics to the mixed Nash equilibrium with respect to the total variation metric. This improves a previous polynomial convergence rate for a similar time-averaged dynamics with different averaging factors. Moreover, unlike previous two-scale approaches for finding the MNE, our approach treats all player types on the same time scale. We also show that with a suitable choice of decreasing temperature, a simulated annealing version of the mean-field dynamics converges to an MNE of the initial unregularized problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。