用深度强化学习解决非平稳连续群体博弈问题
Solving Continuous Mean Field Games: Deep Reinforcement Learning for Non-Stationary Dynamics
- 基于虚构博弈框架,结合强化学习与监督学习求解最优响应
- 通过条件归一化流建模随时间变化的群体分布,提升密度逼近能力
- 在三个复杂场景中验证有效性,适用于真实多智能体系统
群体博弈(MFGs)已成为建模大规模多智能体系统交互的强大框架。尽管强化学习在MFG领域取得进展,现有方法通常局限于有限状态空间或平稳模型,限制了其在现实问题中的应用。本文提出一种专为非平稳连续群体博弈设计的新型深度强化学习算法。该方法基于虚构博弈(FP)框架,利用深度强化学习计算最优响应,并采用监督学习表示平均策略;同时,通过条件归一化流学习随时间变化的群体分布表示。为验证方法有效性,我们在三个复杂度递增的实例上进行了评估。该工作通过解决可扩展性与密度逼近的关键瓶颈,显著推进了深度强化学习在复杂群体博弈问题中的应用,使该领域更接近真实多智能体系统的落地。
原文摘要 · Abstract (English)
Mean field games (MFGs) have emerged as a powerful framework for modeling interactions in large-scale multi-agent systems. Despite recent advancements in reinforcement learning (RL) for MFGs, existing methods are typically limited to finite spaces or stationary models, hindering their applicability to real-world problems. This paper introduces a novel deep reinforcement learning (DRL) algorithm specifically designed for non-stationary continuous MFGs. The proposed approach builds upon a Fictitious Play (FP) methodology, leveraging DRL for best-response computation and supervised learning for average policy representation. Furthermore, it learns a representation of the time-dependent population distribution using a Conditional Normalizing Flow. To validate the effectiveness of our method, we evaluate it on three different examples of increasing complexity. By addressing critical limitations in scalability and density approximation, this work represents a significant advancement in applying DRL techniques to complex MFG problems, bringing the field closer to real-world multi-agent systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。