arXiv:2604.27372math.OCcs.LG2026-04被引 6

提出连续时间平均场控制下的q函数理论,解决带共同噪声的最优策略求解难题。

Continuous-time q-learning for mean-field control with common noise, part-I: Theoretical foundations

  • 基于松弛控制框架推导探索性HJB方程,处理共同噪声带来的非线性策略项
  • 在凹性条件下证明最优单步策略迭代的存在唯一性,支持熵正则化优化
  • 在一般线性二次模型中显式给出最优策略为高斯分布,适用于大规模协同系统

本文研究熵正则化平均场控制(MFC)在受控共同噪声下的连续时间版本,借鉴Jia和Zhou(2023)在单智能体模型中的定义,提出该框架下的q函数。当动作离散采样时,探索性设定下的价值函数随时间网格细化收敛至松弛控制设定下的值函数。基于松弛控制形式,推导出探索性汉密尔顿-雅可比-贝尔曼(HJB)方程,其中受控共同噪声引入策略的额外非线性泛函,使策略迭代复杂化。在特定凹性条件下,通过策略的偏线性泛函导数的一阶条件,建立最优单步策略迭代的存在与唯一性。每次策略改进通过与策略空间上的熵正则化优化问题关联加以验证。在平均场设置下,引入定义于状态分布与策略上的积分q函数(Iq函数),并证明最优策略是Iq函数argmax算子的二层不动点。最后,在一般线性二次(LQ)设定下,显式刻画最优策略为高斯分布。

原文摘要 · Abstract (English)

This paper investigates the continuous-time counterpart of the Q-function for entropy-regularized mean-field control (MFC) with controlled common noise, coined as q-function by Jia and Zhou (2023) in the single agent's model. We first show that, under discretely sampled actions, the value function in the exploratory formulation converges to the one in the relaxed control formulation as the time grid refines. Leveraging the relaxed control formulation, we derive the exploratory Hamilton-Jacobi-Bellman (HJB) equation, in which the controlled common noise gives rise to an additional nonlinear functional of policy, rendering the policy iteration intricate. Under certain concavity condition, we establish the existence and uniqueness of the optimal one-step policy iteration via a first-order condition using the partial linear functional derivative with respect to policy. The policy improvement at each iteration is verified by relating to an entropy-regularized optimization problem over the space of policies. In the mean-field setting, we introduce the integrated q-function (Iq-function) defined on the state distribution and the policy, and it is shown that an optimal policy is identified as a two-layer fixed point to the argmax operator of the Iq-function. Finally, we provide the explicit characterization of an optimal policy as a Gaussian distribution in the general linear-quadratic (LQ) setting.

平均场控制强化学习最优控制连续时间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。