提出新方法认证非单调博弈的收敛性,突破传统限制。
Small-Gain Nash: Certified Contraction to Nash Equilibria in Differentiable Games
- 设计自定义块加权度量,使非单调梯度在局部变为强单调。
- 给出可计算的收敛保证与安全步长,支持单步长动态。
- 适用于复杂博弈,尤其适合带耦合项的强化学习场景。
经典梯度学习收敛性分析依赖于伪梯度在欧氏几何下的(强)单调性,但该条件在存在强跨玩家耦合的简单博弈中常不成立。本文提出小型增益纳什(Small-Gain Nash, SGN),一种基于自定义块加权几何的块小增益条件。SGN将局部曲率与跨玩家利普希茨耦合界转化为可计算的收缩性证书,在满足这些边界时,构造出使伪梯度强单调的加权块度量,即使其在欧氏意义下非单调。连续流在此设计几何中指数收缩,投影欧拉和RK4离散化在显式步长约束下收敛,该约束由SGN裕度和局部利普希茨常数决定。分析揭示了一个非渐近、基于度量的“时间尺度带”,类似TTUR机制:无需通过趋零、不等步长实现渐近时间分离,而是识别一个有限相对度量权重区间,使得单步长动态可证明收缩。我们在二次博弈上验证该框架,欧氏单调性分析无法预测收敛,而SGN成功认证;并扩展至镜像/费舍尔几何,用于马尔可夫博弈中的熵正则化策略梯度。结果是一个离线认证流水线,可在紧集上估计曲率、耦合与利普希茨参数,优化块权重以扩大SGN裕度,并返回结构化、可计算的收敛证书——包含度量、收缩率及非单调博弈的安全步长。
原文摘要 · Abstract (English)
Classical convergence guarantees for gradient-based learning in games require the pseudo-gradient to be (strongly) monotone in Euclidean geometry as shown by rosen(1965), a condition that often fails even in simple games with strong cross-player couplings. We introduce Small-Gain Nash (SGN), a block small-gain condition in a custom block-weighted geometry. SGN converts local curvature and cross-player Lipschitz coupling bounds into a tractable certificate of contraction. It constructs a weighted block metric in which the pseudo-gradient becomes strongly monotone on any region where these bounds hold, even when it is non-monotone in the Euclidean sense. The continuous flow is exponentially contracting in this designed geometry, and projected Euler and RK4 discretizations converge under explicit step-size bounds derived from the SGN margin and a local Lipschitz constant. Our analysis reveals a certified "timescale band", a non-asymptotic, metric-based certificate that plays a TTUR-like role: rather than forcing asymptotic timescale separation via vanishing, unequal step sizes, SGN identifies a finite band of relative metric weights for which a single-step-size dynamics is provably contractive. We validate the framework on quadratic games where Euclidean monotonicity analysis fails to predict convergence, but SGN successfully certifies it, and extend the construction to mirror/Fisher geometries for entropy-regularized policy gradient in Markov games. The result is an offline certification pipeline that estimates curvature, coupling, and Lipschitz parameters on compact regions, optimizes block weights to enlarge the SGN margin, and returns a structural, computable convergence certificate consisting of a metric, contraction rate, and safe step-sizes for non-monotone games.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。