从离散数据中学习连续时间群体智能控制,精度可达Δt阶。
Mean-Field PhiBE: Continuous-Time Mean-Field Reinforcement Learning from Discrete-Time Data

- 用一步数据估计连续方程中的漂移与扩散项,保持微分结构。
- 在线性二次情形下仅用一步数据即达二阶精度,误差为Δt²。
- 适合研究群体行为建模、强化学习与连续时间控制的科研人员。
本文研究模型无关的连续时间平均场控制问题:群体动态由未知的麦凯恩-弗拉索夫随机微分方程连续演化,但仅可获取离散时间转移数据。在模型依赖框架下,策略评估自然由定义在概率测度空间 $\mathcal P_2(\mathbb R^d)$ 上的平稳哈密顿-雅可比-贝尔曼方程描述,但该方程涉及受控麦凯恩-弗拉索夫动态的漂移与扩散系数,在仅有离散数据时不可识别。而直接降维至离散贝尔曼方程虽避免不可识别性,却丧失微分方程结构。为此,我们提出平均场-Φ贝尔曼方程(MF-PhiBE),将离散时间转移信息嵌入沃尔瑟斯坦空间上的连续时间偏微分方程。MF-PhiBE用数据计算的一步估计量替代未知的无穷小漂移与协方差,同时保留麦凯恩-弗拉索夫哈密顿-雅可比-贝尔曼方程的生成器结构。我们还推导了熵正则化随机反馈策略的策略梯度定理,通过动作级无穷小优势与策略得分表达演员方向。结合二者,得到一个模型无关的演员-评论家方法。我们证明一阶一致性估计:最优MF-PhiBE策略诱导的值函数与最优连续时间值函数之间的误差为 $Δt$ 阶。在线性二次情形下,我们证明该近似仅用一步数据即可达到二阶精度。在LQR基准和人群回避问题上的数值实验验证了该框架的有效性。
原文摘要 · Abstract (English)
This paper addresses model-free continuous-time mean-field control in a setting where the population dynamics evolve continuously according to an unknown McKean-Vlasov stochastic differential equation, while only discrete-time transition data are available. In the model-based formulation, policy evaluation is naturally described by a stationary Hamilton-Jacobi-Bellman equation on $\mathcal P_2(\mathbb R^d)$, but this equation involves the drift and diffusion coefficients of the controlled McKean-Vlasov dynamics, which are not identifiable when only discrete-time data are available. On the other hand, a direct reduction to a time-discrete Bellman equation avoids the non-identifiability issue but loses the differential equation structure. To bridge these two viewpoints, we introduce a Mean-Field-PhiBE (MF-PhiBE), which incorporates discrete-time transition information into a continuous-time PDE on the Wasserstein space. The MF-PhiBE replaces the unknown infinitesimal drift and covariance in the policy-evaluation equation by one-step estimators computed from data, while preserving the generator structure of the McKean-Vlasov HJB equation. We also derive a policy-gradient theorem for entropy-regularized randomized feedback policies, expressing the actor direction through an action-wise infinitesimal advantage and the score of the policy. Combining these two ingredients yields a model-free actor-critic method. We prove a first-order consistency estimate showing that the value induced by an optimal MF-PhiBE policy approximates the optimal continuous-time value with an error of order $Δt$. In the linear-quadratic case, we show our approximation achieves second-order accuracy with only one-step data. Numerical experiments on an LQR benchmark and a crowd-aversion problem illustrate the proposed framework.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。