arXiv:2601.11217math.OCcs.LG2026-01被引 2

提出无需环境模型的均值场控制策略梯度算法,解决群体状态依赖难题。

Model-free policy gradient for discrete-time mean-field control

  • 通过扰动状态分布流构造新型梯度估计器
  • 实现完全模型无关的策略优化,误差有定量保证
  • 适合大规模多智能体系统协同控制研究者

我们研究有限状态空间与紧致动作空间下离散时间均值场控制(MFC)问题的无模型策略学习。相较于已有大量基于价值的方法,由于转移核与奖励函数依赖于随时间演化的群体状态分布,传统单智能体策略梯度的似然比估计器难以直接应用,导致基于策略的方法长期未被充分探索。本文提出一种新的状态分布流扰动方案,证明扰动后的价值函数梯度在扰动强度趋于零时收敛至真实策略梯度。该构造催生了一个仅依赖模拟轨迹和状态分布敏感性辅助估计的全模型无关梯度估计器。在此框架基础上,我们设计了MF-REINFORCE算法,并给出了其偏差与均方误差的显式量化边界。在典型均值场控制任务上的数值实验验证了该方法的有效性。

原文摘要 · Abstract (English)

We study model-free policy learning for discrete-time mean-field control (MFC) problems with finite state space and compact action space. In contrast to the extensive literature on value-based methods for MFC, policy-based approaches remain largely unexplored due to the intrinsic dependence of transition kernels and rewards on the evolving population state distribution, which prevents the direct use of likelihood-ratio estimators of policy gradients from classical single-agent reinforcement learning. We introduce a novel perturbation scheme on the state-distribution flow and prove that the gradient of the resulting perturbed value function converges to the true policy gradient as the perturbation magnitude vanishes. This construction yields a fully model-free estimator based solely on simulated trajectories and an auxiliary estimate of the sensitivity of the state distribution. Building on this framework, we develop MF-REINFORCE, a model-free policy gradient algorithm for MFC, and establish explicit quantitative bounds on its bias and mean-squared error. Numerical experiments on representative mean-field control tasks demonstrate the effectiveness of the proposed approach.

强化学习多智能体策略梯度均值场控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。