arXiv:2510.12180math.OCcs.LG2025-10被引 2

用连续动态求解大规模博弈,让智能体分布自然收敛到均衡。

Learning Mean-Field Games through Mean-Field Actor-Critic Flow

  • 通过偏微分方程联合更新策略、价值函数和分布
  • 在合适时间尺度下实现全局指数收敛
  • 适合研究群体智能与大规模博弈的算法设计

我们提出均值场演员-评论家(MFAC)流,一种用于求解均值场博弈(MFGs)的连续时间学习动力学,结合了强化学习与最优传输技术。MFAC 框架通过由偏微分方程(PDEs)控制的耦合梯度更新,同步演化控制策略(演员)、价值函数(评论家)和分布组件。核心创新是最优传输测地线皮卡德(OTGP)流,它沿 Wasserstein-2 测地线驱动分布向均衡收敛。我们利用李雅普诺夫泛函进行了严格的收敛性分析,证明在合适时间尺度下 MFAC 流具有全局指数收敛性。结果揭示了演员、评论家与分布组件之间的算法协同作用。数值实验验证了理论发现,并展示了 MFAC 框架在计算 MFG 均衡方面的有效性。

原文摘要 · Abstract (English)

We propose the Mean-Field Actor-Critic (MFAC) flow, a continuous-time learning dynamics for solving mean-field games (MFGs), combining techniques from reinforcement learning and optimal transport. The MFAC framework jointly evolves the control (actor), value function (critic), and distribution components through coupled gradient-based updates governed by partial differential equations (PDEs). A central innovation is the Optimal Transport Geodesic Picard (OTGP) flow, which drives the distribution toward equilibrium along Wasserstein-2 geodesics. We conduct a rigorous convergence analysis using Lyapunov functionals and establish global exponential convergence of the MFAC flow under a suitable timescale. Our results highlight the algorithmic interplay among actor, critic, and distribution components. Numerical experiments illustrate the theoretical findings and demonstrate the effectiveness of the MFAC framework in computing MFG equilibria.

博弈论强化学习最优传输连续动力学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。