arXiv:2511.06812math.OCcs.LG2025-11被引 3

证明了连续空间下演员-评论家算法在平均场问题中的收敛性。

Convergence of Actor-Critic Learning for Mean Field Games and Mean Field Control in Continuous Spaces

  • 基于双时间尺度框架,统一处理平均场博弈与控制问题。
  • 在一二维线性二次模型中验证算法收敛且结果逼近理论解。
  • 适用于群体局部合作全局竞争的复杂场景,具理论严谨性。

我们建立了[Angiuli等, 2023a]提出的深度演员-评论家强化学习算法在连续状态和动作空间、无限时域下的收敛性。该算法根据价值函数与平均场项学习率的比率,可求解平均场博弈(MFG)或平均场控制(MFC)问题。在MFC情形中,为严格识别极限,我们采用[Angiuli等, 2023b]在有限空间中的方法对状态和动作空间进行离散化。收敛性证明基于[Borkar, 1997]提出的双时间尺度框架的推广。我们进一步将结果扩展至包含局部合作与全局竞争群体的平均场控制博弈。最后,我们在一维和二维线性二次问题上进行了数值实验,这些情况下存在显式解,用于验证算法性能。

原文摘要 · Abstract (English)

We establish the convergence of the deep actor-critic reinforcement learning algorithm presented in [Angiuli et al., 2023a] in the setting of continuous state and action spaces with an infinite discrete-time horizon. This algorithm provides solutions to Mean Field Game (MFG) or Mean Field Control (MFC) problems depending on the ratio between two learning rates: one for the value function and the other for the mean field term. In the MFC case, to rigorously identify the limit, we introduce a discretization of the state and action spaces, following the approach used in the finite-space case in [Angiuli et al., 2023b]. The convergence proofs rely on a generalization of the two-timescale framework introduced in [Borkar, 1997]. We further extend our convergence results to Mean Field Control Games, which involve locally cooperative and globally competitive populations. Finally, we present numerical experiments for linear-quadratic problems in one and two dimensions, for which explicit solutions are available.

强化学习平均场收敛性控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。