arXiv:2504.20593cs.LG2025-04被引 4

研究多智能体在动态环境中的独立学习,发现稳定解存在且算法可收敛。

Independent Learning in Performative Markov Potential Games

  • 引入可实现稳定均衡概念,证明其在合理假设下恒存在。
  • 独立策略梯度算法在近似意义下收敛到稳定解,性能受可实现效应影响。
  • 适用于智能体独立优化、环境动态变化的场景,如市场博弈或资源分配。

表现性强化学习(PRL)指部署的策略会改变环境的奖励与转移动态。本文将表现性效应引入马尔可夫势博弈(MPG),提出可实现稳定均衡(PSE)概念,并在合理敏感性假设下证明其必然存在。我们分析了当前主流算法在解决MPG时的收敛性:独立策略梯度上升(IPGA)和独立自然策略梯度(INPG)在最优迭代意义下收敛至近似PSE,额外项反映表现性效应影响;此外,INPG在最后迭代意义下渐进收敛至PSE。当表现性效应消失时,恢复已有收敛速率。针对特定情形,我们还提供了重复再训练方法的有限时间最后迭代收敛结果,其中智能体独立优化代理目标。通过大量实验验证了理论结论。

原文摘要 · Abstract (English)

Performative Reinforcement Learning (PRL) refers to a scenario in which the deployed policy changes the reward and transition dynamics of the underlying environment. In this work, we study multi-agent PRL by incorporating performative effects into Markov Potential Games (MPGs). We introduce the notion of a performatively stable equilibrium (PSE) and show that it always exists under a reasonable sensitivity assumption. We then provide convergence results for state-of-the-art algorithms used to solve MPGs. Specifically, we show that independent policy gradient ascent (IPGA) and independent natural policy gradient (INPG) converge to an approximate PSE in the best-iterate sense, with an additional term that accounts for the performative effects. Furthermore, we show that INPG asymptotically converges to a PSE in the last-iterate sense. As the performative effects vanish, we recover the convergence rates from prior work. For a special case of our game, we provide finite-time last-iterate convergence results for a repeated retraining approach, in which agents independently optimize a surrogate objective. We conduct extensive experiments to validate our theoretical findings.

多智能体强化学习博弈论收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。