arXiv:2501.00052cs.LGcs.GT2025-01

提出高效可扩展的深度强化学习方法求解大规模群体控制问题

Efficient and Scalable Deep Reinforcement Learning for Mean Field Control Games

  • 将无限智能体系统转为马尔可夫决策过程,用演员-评论家框架并行优化
  • 在线性二次模型上实现样本效率提升一个数量级,逼近理论最优解
  • 适合研究交通流、经济竞争等大规模多智能体系统建模与优化

均值场控制博弈(MFCG)为分析无限多个交互智能体的系统提供了有力的理论框架,融合了均值场博弈(MFG)和均值场控制(MFC)的要素。然而,求解刻画MFCG均衡的耦合汉密尔顿-雅可比-贝尔曼方程与福克-普朗克方程仍是重大计算挑战,尤其在高维或复杂环境中。本文提出一种可扩展的深度强化学习(RL)方法,用于近似MFCG的均衡解。基于前期工作,我们将无限智能体随机控制问题重构为马尔可夫决策过程,每个代表性智能体与随时间演化的均值场分布交互。以Angiuli等人(2024)提出的演员-评论家算法为基础,我们设计了多种更高效、更可扩展的算法版本,采用并行采样(批处理)、小批量训练、目标网络、近端策略优化(PPO)、广义优势估计(GAE)和熵正则化等技术。通过这些改进,显著提升了基线算法的效率、可扩展性和训练稳定性。我们在一个线性二次基准问题上评估方法,该问题存在解析解。结果表明,部分所提方法实现更快收敛,且接近理论最优,样本效率相比基线提升一个数量级。本工作为将深度强化学习应用于更复杂的现实相关MFCG(如大规模自动驾驶交通系统、多企业经济竞争、银行间借贷问题)奠定了基础。

原文摘要 · Abstract (English)

Mean Field Control Games (MFCGs) provide a powerful theoretical framework for analyzing systems of infinitely many interacting agents, blending elements from Mean Field Games (MFGs) and Mean Field Control (MFC). However, solving the coupled Hamilton-Jacobi-Bellman and Fokker-Planck equations that characterize MFCG equilibria remains a significant computational challenge, particularly in high-dimensional or complex environments. This paper presents a scalable deep Reinforcement Learning (RL) approach to approximate equilibrium solutions of MFCGs. Building on previous works, We reformulate the infinite-agent stochastic control problem as a Markov Decision Process, where each representative agent interacts with the evolving mean field distribution. We use the actor-critic based algorithm from a previous paper (Angiuli et.al., 2024) as the baseline and propose several versions of more scalable and efficient algorithms, utilizing techniques including parallel sample collection (batching); mini-batching; target network; proximal policy optimization (PPO); generalized advantage estimation (GAE); and entropy regularization. By leveraging these techniques, we effectively improved the efficiency, scalability, and training stability of the baseline algorithm. We evaluate our method on a linear-quadratic benchmark problem, where an analytical solution to the MFCG equilibrium is available. Our results show that some versions of our proposed approach achieve faster convergence and closely approximate the theoretical optimum, outperforming the baseline algorithm by an order of magnitude in sample efficiency. Our work lays the foundation for adapting deep RL to solve more complicated MFCGs closely related to real life, such as large-scale autonomous transportation systems, multi-firm economic competition, and inter-bank borrowing problems.

强化学习多智能体均值场优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。