研究大规模群体智能体的平均场强化学习,打通控制理论与强化学习的桥梁。
Mean Field Reinforcement Learning

- 从多智能体系统出发,构建平均场交互下的马尔可夫决策模型
- 证明了有限群体与无限群体系统的渐近一致性,支持理论分析
- 适合研究大规模智能体系统的学者,尤其关注数学建模与算法设计
本专著从大群体随机控制中具有平均场交互和公共噪声的马尔可夫决策过程视角,介绍平均场强化学习。从多智能体强化学习与平均场控制的联系出发,构建表征性智能体学习问题的概率、数学与控制论框架,分析其与有限群体系统的关联,并研究一般及线性-二次模型。内容涵盖动态规划原理、混沌传播极限,以及表格型Q-learning和策略梯度方法的理论分析。还讨论数值实现,包括表格方法与深度强化学习如深度确定性策略梯度。目标是为平均场控制理论与强化学习方法之间建立连贯桥梁,强调问题的数学结构与对大规模随机群体的可计算学习方法设计。
原文摘要 · Abstract (English)
This monograph provides an introduction to mean field reinforcement learning through the lens of Markov decision processes arising from large-population stochastic control with mean field interactions and common noise. Starting from the connection between multi-agent reinforcement learning and mean field control, it develops the probabilistic, mathematical, and control-theoretic framework needed to formulate representative-agent learning problems, analyze their relationship with finite-population systems, and study both general and linear-quadratic models. The presentation includes dynamic programming principles, propagation-of-chaos limits, and theoretical analyses of tabular Q-learning and policy-gradient methods. It also discusses numerical implementations, including tabular schemes and deep reinforcement learning methods such as deep deterministic policy gradient. The goal is to give readers a coherent bridge between mean field control theory and reinforcement learning methodology, emphasizing the mathematical structure of the problems and the design of tractable learning approaches for large stochastic populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。