arXiv:2605.13025cs.LGcs.GT2026-05被引 4

仅用KL正则化就能稳定学习并收敛到纳什均衡,无需显式悲观策略。

Offline Two-Player Zero-Sum Markov Games with KL Regularization

  • 用KL正则化替代传统悲观性,构建理论框架ROSE。
  • 实现$ ilde{ m O}(1/n)$的收敛速度,优于未正则化时的$ ilde{ m O}(1/ ext{√}n)$。
  • 提出SOS-MD算法,适合无模型、离线的双人零和博弈学习。

我们研究离线双人零和马尔可夫博弈中纳什均衡的学习问题。现有方法常依赖显式悲观性应对分布偏移,而我们证明仅用KL正则化即可稳定学习并保证收敛。首先提出正则化离线序贯均衡(ROSE)理论框架,在单边集中性条件下达到$ ilde{ m O}(1/n)$的快速收敛率,优于未正则化设置下的标准$ ilde{ m O}(1/ ext{√}n)$。随后提出基于最小二乘值估计与迭代自对弈更新的实用无模型算法SOS-MD。证明其最后迭代在自对弈轮次$T$下,以$ ilde{ m O}(1/ ext{√}T)$的优化误差达成相同的$ ilde{ m O}(1/n)$统计速率。

原文摘要 · Abstract (English)

We study the problem of learning Nash equilibria in offline two-player zero-sum Markov games. While existing approaches often rely on explicit pessimism to address distribution shift, we show that KL regularization alone suffices to stabilize learning and guarantee convergence. We first introduce Regularized Offline Sequential Equilibrium (ROSE), a theoretical framework that achieves a fast $\widetilde{\mathcal{O}}(1/n)$ convergence rate under \textit{unilateral concentrability}, improving over the standard $\widetilde{\mathcal{O}}(1/\sqrt{n})$ rates in unregularized settings. We then propose Sequential Offline Self-play Mirror Descent (SOS-MD), a practical model-free algorithm based on least-squares value estimation and iterative self-play updates. We prove that the last iterate of SOS-MD attains the same $\widetilde{\mathcal{O}}(1/n)$ statistical rate up to a vanishing optimization error of order $\widetilde{\mathcal{O}}(1/\sqrt{T})$ in the number of self-play iterations $T$.

博弈学习离线强化学习纳什均衡正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。