arXiv:2604.27378math.OCcs.LG2026-04被引 5

提出适用于带共同噪声均场控制的连续时间Q-learning算法,实现可学习的最优策略更新。

Continuous-time q-learning for mean-field control with common noise, part-II: q-learning algorithms

  • 基于松弛控制框架,通过条件状态分布验证值函数与Iq函数的鞅性质。
  • 在离散采样动作下,量化了不可观测数据替换带来的误差,保证算法可行性。
  • 设计了演员-评论家型Q-learning算法,在线性二次和非线性场景中表现良好。

本文是Ren等(2026)工作的延续,旨在为带有受控共同噪声的均场控制(MFC)进一步设计Q-learning算法。基于松弛控制框架,我们首先通过在所有测试策略生成的条件状态分布上评估,建立了值函数与Iq函数的鞅条件。由于松弛控制中的数据在实际中不可观测,我们量化了在离散采样动作下,用可观察的探索性公式数据替代时引入的误差。结合Ren等(2026)提出的最优策略双层不动点表征,我们提出了多种算法,包括演员-评论家型Q-learning算法:演员步基于改进的Iq函数迭代规则更新策略,评论家步则利用探索性公式数据,依据鞅正交性条件更新值函数与Iq函数。我们在无限时域线性二次(LQ)框架下建立了演员步内迭代的收敛性。在两个例子中——一个在LQ框架内,一个超出该框架——我们的算法实现了满意的性能。

原文摘要 · Abstract (English)

This paper is a continuation work of Ren et al. (2026) aiming to further devise q-learning algorithms for mean-field control (MFC) with controlled common noise. Based on the relaxed control formulation, we first establish the martingale condition of the value function and the Iq-function by evaluating along the conditional state distributions generated by all test policies. As the data in the relaxed control formulation are not observable in practice, we quantify the error incurred when they are replaced by the observable ones in the exploratory formulation under discretely sampled actions. This, together with a two-layer fixed point characterization of an optimal policy in Ren et al. (2026), allows us to propose several algorithms including the Actor-Critic q-learning algorithm, in which the policy is updated in the Actor-step based on the iteration rule induced by the improved Iq-function, and the value function and Iq-function are updated in the Critic-step based on the martingale orthogonality condition using the data from the exploratory formulation. We also establish the convergence of the inner iterations in the Actor-step in an infinite-horizon linear quadratic (LQ) framework. In two examples, within and beyond LQ framework, our q-learning algorithms are implemented with satisfactory performance.

强化学习均场控制连续时间Q-learning

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。