arXiv:2412.20203cs.GTcs.LG2024-12NeurIPS被引 9

研究多智能体学习在利益冲突场景下的收敛性,发现改进算法可稳定达成均衡。

No-regret learning in harmonic games: Extrapolation in the face of conflicting interests

  • 引入外推机制改进FTRL算法,提升学习稳定性。
  • 在任意初始条件下,算法可收敛至纳什均衡,且累积损失不超过常数。
  • 揭示了冲突博弈与协同博弈在动态行为上的对称关系,适合博弈论与强化学习研究者。

多智能体学习的长期行为,尤其是无遗憾学习,在势博弈(玩家利益一致)中已有较好理解。然而,在谐振博弈(玩家利益冲突的战略对应物)中,除双人零和博弈且具有完全混合均衡这一狭窄子类外,几乎一无所知。本文聚焦于广义谐振博弈,分析最广泛研究的无遗憾学习方案——跟随正则化领导者(FTRL)的收敛性质。首先,我们证明连续时间下FTRL动力学具有庞加莱回归性,即无限次返回初始点附近,因而无法收敛。离散时间下,标准的FTRL实现可能导致更差结果,使玩家陷入持续的最佳回应循环。但若在FTRL中加入合适的外推步骤(包括乐观和镜像逼近变体),我们证明学习可从任意初始条件收敛至纳什均衡,且所有玩家的遗憾值均保证为O(1)。这些结果深入揭示了谐振博弈中无遗憾学习的动态特性,统一了先前对双人零和博弈的研究,并表明从战略与动态双重角度看,谐振博弈是势博弈的自然互补。

原文摘要 · Abstract (English)

The long-run behavior of multi-agent learning - and, in particular, no-regret learning - is relatively well-understood in potential games, where players have aligned interests. By contrast, in harmonic games - the strategic counterpart of potential games, where players have conflicting interests - very little is known outside the narrow subclass of 2-player zero-sum games with a fully-mixed equilibrium. Our paper seeks to partially fill this gap by focusing on the full class of (generalized) harmonic games and examining the convergence properties of follow-the-regularized-leader (FTRL), the most widely studied class of no-regret learning schemes. As a first result, we show that the continuous-time dynamics of FTRL are Poincaré recurrent, that is, they return arbitrarily close to their starting point infinitely often, and hence fail to converge. In discrete time, the standard, "vanilla" implementation of FTRL may lead to even worse outcomes, eventually trapping the players in a perpetual cycle of best-responses. However, if FTRL is augmented with a suitable extrapolation step - which includes as special cases the optimistic and mirror-prox variants of FTRL - we show that learning converges to a Nash equilibrium from any initial condition, and all players are guaranteed at most O(1) regret. These results provide an in-depth understanding of no-regret learning in harmonic games, nesting prior work on 2-player zero-sum games, and showing at a high level that harmonic games are the canonical complement of potential games, not only from a strategic, but also from a dynamic viewpoint.

博弈论多智能体无遗憾学习纳什均衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。