用强化学习求解大规模群体博弈的纳什均衡,方法可扩展且高效。
Reinforcement Learning for Finite Space Mean-Field Type Games
- 基于均值场空间量化与纳什Q学习,设计可收敛的算法
- 在4个环境中验证,支持维度达200的均值场分布
- 适合研究大规模多智能体协作与竞争的学者
均值场类型博弈(MFTGs)描述了由大量群体组成的纳什均衡:每个群体由连续数量的合作者构成,最大化其群体平均收益,同时与其他有限数量的群体非合作互动。尽管理论已较为完善,但高效且可扩展的计算方法仍不足。本文在一般动态和奖励函数的有限空间设定下,发展了强化学习方法。首先证明,MFTG解可近似为有限规模群体博弈的纳什均衡。随后提出两种算法:第一种基于均值场空间量化与纳什Q学习,给出收敛性与稳定性分析;第二种为深度强化学习算法,可扩展至更大空间。在4个环境中的数值实验显示,该方法在均值场分布维度高达200时仍具可扩展性与效率。
原文摘要 · Abstract (English)
Mean field type games (MFTGs) describe Nash equilibria between large coalitions: each coalition consists of a continuum of cooperative agents who maximize the average reward of their coalition while interacting non-cooperatively with a finite number of other coalitions. Although the theory has been extensively developed, we are still lacking efficient and scalable computational methods. Here, we develop reinforcement learning methods for such games in a finite space setting with general dynamics and reward functions. We start by proving that the MFTG solution yields approximate Nash equilibria in finite-size coalition games. We then propose two algorithms. The first is based on the quantization of mean-field spaces and Nash Q-learning. We provide convergence and stability analysis. We then propose a deep reinforcement learning algorithm, which can scale to larger spaces. Numerical experiments in 4 environments with mean-field distributions of dimension up to $200$ show the scalability and efficiency of the proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。