arXiv:2506.16120cs.GTcs.LG2025-06ICML被引 4

首次证明独立策略梯度可收敛至零和凸马尔可夫博弈的纳什均衡。

Solving Zero-Sum Convex Markov Games

  • 通过非凸正则化将问题转化为NC-pPL目标,稳定迭代过程。
  • 在无限时域下实现随机嵌套与交替梯度方法的全局收敛。
  • 适用于多智能体战略交互建模,对强化学习研究者有重要参考价值。

我们首次为两玩家零和凸马尔可夫博弈(cMGs)提供了独立策略梯度方法全局收敛至纳什均衡的可证明保证。凸马尔可夫博弈由Gemp等(2024)提出,是将马尔可夫决策过程推广到多智能体场景的框架,其状态偏好在占用测度上呈凸性,能广泛建模一般战略互动。然而,即使最基础的极小极大情形也面临严峻挑战:固有的非凸性、贝尔曼一致性缺失以及无限时域复杂性。本文采用两步法:首先,利用隐藏凸-隐藏凹函数性质,证明简单非凸正则化可将极小极大优化转化为非凸-近端Polyak-Lojasiewicz(NC-pPL)目标;关键在于该正则化能稳定独立策略梯度的迭代,最终引导其收敛至均衡。其次,在此基础上,针对满足NC-pPL与双侧pPL条件的广义约束极小极大问题,首次提供随机嵌套与交替梯度下降-上升方法的全局收敛保证,这些结果可能具有独立研究价值。

原文摘要 · Abstract (English)

We contribute the first provable guarantees of global convergence to Nash equilibria (NE) in two-player zero-sum convex Markov games (cMGs) by using independent policy gradient methods. Convex Markov games, recently defined by Gemp et al. (2024), extend Markov decision processes to multi-agent settings with preferences that are convex over occupancy measures, offering a broad framework for modeling generic strategic interactions. However, even the fundamental min-max case of cMGs presents significant challenges, including inherent nonconvexity, the absence of Bellman consistency, and the complexity of the infinite horizon. We follow a two-step approach. First, leveraging properties of hidden-convex--hidden-concave functions, we show that a simple nonconvex regularization transforms the min-max optimization problem into a nonconvex-proximal Polyak-Lojasiewicz (NC-pPL) objective. Crucially, this regularization can stabilize the iterates of independent policy gradient methods and ultimately lead them to converge to equilibria. Second, building on this reduction, we address the general constrained min-max problems under NC-pPL and two-sided pPL conditions, providing the first global convergence guarantees for stochastic nested and alternating gradient descent-ascent methods, which we believe may be of independent interest.

博弈论强化学习收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。