arXiv:2607.11005math.OCcs.AI2026-07

提出一种连续时间均场控制的确定性策略强化学习方法。

Actor-Critic Learning for Extended Mean Field Control with Deterministic Policies

  • 采用确定性反馈策略,避免随机核优化难题。
  • 在Cucker-Smale共识与交易拥挤最优清算中验证高效稳定。
  • 适合处理状态-控制联合分布依赖的复杂系统建模。

本文构建了一种针对连续时间扩展均场控制问题的无模型强化学习框架,其中动态和奖励均依赖于状态与控制的联合分布。采用确定性反馈策略,使状态-动作分布直接由状态分布通过前推映射生成,避免了对随机核的优化,突破了现有扩展均场方法的关键局限。首先建立了参数化麦克斯韦-弗拉索夫动力学的无模型敏感性公式,并导出基于沃瑟斯坦空间优势率函数的确定性策略梯度表达式。随后通过引入依赖于状态、动作及联合状态-动作分布的局部值与优势率表示,得到包含动作导数和控制分布测度导数项的策略梯度。这些表征催生了基于鞅的学习原则,并启发了一种结合粒子近似、测度依赖神经网络、时序差分学习及动作或参数空间探索的连续时间深度确定性策略梯度算法。在随机Cucker-Smale共识控制和具有交易拥挤效应的最优清仓问题上的数值实验表明,该方法具备高效性、稳定性与鲁棒性,包括显式依赖控制分布的问题。

原文摘要 · Abstract (English)

This paper develops a model-free reinforcement learning framework for continuous--time extended mean field control problems, where both the dynamics and reward may depend on the joint distribution of states and controls. We adopt deterministic feedback policies, under which the state--action distribution is induced directly as a push--forward of the state law. This avoids optimization over stochastic kernels and bypasses key limitations of existing approaches in extended mean field settings. We first establish a model--free sensitivity formula for parameterized McKean--Vlasov dynamics and use it to derive a deterministic policy gradient formula expressed through an advantage--rate function on the Wasserstein space. We then refine this formula by introducing local value and advantage--rate representations that depend on the state, action, and joint state--action distribution, yielding a policy gradient that includes both action derivatives and measure--derivative terms with respect to the control distribution. These characterizations lead to a martingale--based learning principle and motivate a continuous--time deep deterministic policy gradient algorithm combining particle approximations, measure--dependent neural networks, temporal--difference learning, and exploration in either action or parameter space. Numerical experiments on stochastic Cucker--Smale consensus control and optimal liquidation with trade crowding demonstrate the efficiency, stability, and robustness of the proposed method, including problems with explicit dependence on the control distribution.

强化学习均场控制确定性策略扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。