arXiv:2409.04613cs.MAcs.AI2024-09被引 6

提出新型分布式算法,让多智能体在复杂博弈中稳定学习。

Convergence of Decentralized Actor-Critic Algorithm in General-sum Markov Games

  • 各智能体用异步更新的演员-评论家策略,独立决策
  • 首次证明一般和博弈下算法收敛到近似均衡策略集
  • 适合研究真实场景中的多智能体协作与竞争

马尔可夫博弈为动态环境中多智能体战略交互提供了强大框架。传统上,去中心化学习算法的收敛性仅在零和博弈和势博弈等特例中被证明,难以刻画真实世界互动。本文填补这一空白,研究一般和博弈中学习算法的渐近性质。重点分析一种去中心化算法,其中每个智能体采用具有异步步长的演员-评论家学习动态。该方法使智能体能独立运作,无需知晓他人策略或收益。我们引入马尔可夫近势函数(MNPF),并证明其作为策略更新的近似李雅普诺夫函数,从而刻画了收敛策略集。在特定正则性条件下且纳什均衡有限时,进一步强化了结果。

原文摘要 · Abstract (English)

Markov games provide a powerful framework for modeling strategic multi-agent interactions in dynamic environments. Traditionally, convergence properties of decentralized learning algorithms in these settings have been established only for special cases, such as Markov zero-sum and potential games, which do not fully capture real-world interactions. In this paper, we address this gap by studying the asymptotic properties of learning algorithms in general-sum Markov games. In particular, we focus on a decentralized algorithm where each agent adopts an actor-critic learning dynamic with asynchronous step sizes. This decentralized approach enables agents to operate independently, without requiring knowledge of others' strategies or payoffs. We introduce the concept of a Markov Near-Potential Function (MNPF) and demonstrate that it serves as an approximate Lyapunov function for the policy updates in the decentralized learning dynamics, which allows us to characterize the convergent set of strategies. We further strengthen our result under specific regularity conditions and with finite Nash equilibria.

多智能体博弈学习收敛分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。