arXiv:2503.06037cs.MAcs.AI2025-03

将强化学习转化为博弈论中的推断问题,实现多智能体博弈的稳定求解。

Vairiational Stochastic Games

  • 基于变分推断构建去中心化多智能体框架,解决目标不一致与非平稳性问题。
  • 证明所提策略可逼近ε-Nash均衡,且算法具备理论收敛性保证。
  • 适用于博弈均衡求解,适合研究多智能体协作与竞争的学者。

控制即推断(CAI)框架已成功将单智能体强化学习转化为概率推断问题。然而,该框架在去中心化多智能体广义随机博弈(SGs)中的扩展仍不充分,尤其在缺乏中央协调的情况下。本文提出一种专为去中心化多智能体系统设计的新型变分推断框架,有效应对非平稳性和目标不一致带来的挑战,并证明所得策略构成ε-Nash均衡。此外,我们还给出了所提去中心化算法的理论收敛性保证。基于该框架,我们实例化多个算法以求解纳什均衡、均值场纳什均衡及相关均衡,并提供严格的理论收敛分析。

原文摘要 · Abstract (English)

The Control as Inference (CAI) framework has successfully transformed single-agent reinforcement learning (RL) by reframing control tasks as probabilistic inference problems. However, the extension of CAI to multi-agent, general-sum stochastic games (SGs) remains underexplored, particularly in decentralized settings where agents operate independently without centralized coordination. In this paper, we propose a novel variational inference framework tailored to decentralized multi-agent systems. Our framework addresses the challenges posed by non-stationarity and unaligned agent objectives, proving that the resulting policies form an $ε$-Nash equilibrium. Additionally, we demonstrate theoretical convergence guarantees for the proposed decentralized algorithms. Leveraging this framework, we instantiate multiple algorithms to solve for Nash equilibrium, mean-field Nash equilibrium, and correlated equilibrium, with rigorous theoretical convergence analysis.

多智能体博弈论变分推断强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。