arXiv:2503.10186cs.MAcs.AI2025-03被引 1

研究多智能体Q学习在随机网络中的收敛性,发现可控交互可实现稳定学习。

Convergence and Connectivity: Dynamics of Multi-Agent Q-Learning in Random Networks

  • 基于随机图模型分析多智能体Q学习动态,关注网络连通性影响。
  • 在特定探索率与交互概率下,系统能收敛至唯一均衡点。
  • 适用于大规模分布式系统设计,尤其适合有社区结构的场景。

除了特定设置外,许多多智能体学习算法无法收敛到均衡解,反而表现出复杂的非平稳行为,如周期或混沌轨道。近期研究表明,随着智能体数量增加,此类复杂行为更可能发生。本文研究在随机网络中的网络多项式正则型博弈中Q学习的动力学特性,重点关注经典随机图模型:用于分析分布式系统连通性的Erdős-Rényi模型,以及考虑社区结构的Stochastic Block模型。在每种设定下,我们建立了代理联合策略收敛至唯一均衡的充分条件。研究揭示该条件依赖于探索率、收益矩阵,以及智能体间交互概率。通过数值模拟验证理论结果,表明只要控制网络中的交互,即可在多智能体系统中可靠实现收敛。

原文摘要 · Abstract (English)

Beyond specific settings, many multi-agent learning algorithms fail to converge to an equilibrium solution, instead displaying complex, non-stationary behaviours such as recurrent or chaotic orbits. In fact, recent literature suggests that such complex behaviours are likely to occur when the number of agents increases. In this paper, we study Q-learning dynamics in network polymatrix normal-form games where the network structure is drawn from classical random graph models. In particular, we focus on the Erdős-Rényi model, which is used to analyze connectivity in distributed systems, and the Stochastic Block model, which generalizes the above by accounting for community structures that naturally arise in multi-agent systems. In each setting, we establish sufficient conditions under which the agents' joint strategies converge to a unique equilibrium. We investigate how this condition depends on the exploration rates, payoff matrices and, crucially, the probabilities of interaction between network agents. We validate our theoretical findings through numerical simulations and demonstrate that convergence can be reliably achieved in many-agent systems, provided interactions in the network are controlled.

多智能体Q学习随机网络收敛性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。