arXiv:2509.15392cs.LG2025-09NeurIPS被引 1

提出首个非渐近收敛的领导者-追随者群体博弈学习算法

Learning in Stackelberg Mean Field Games: A Non-Asymptotic Analysis

  • 单循环架构交替更新领导者、代表性跟随者和群体均值策略
  • 在有限步内收敛至稳定点,无需领导者与跟随者独立的强假设
  • 适用于经济建模等大规模博弈场景,优于现有多智能体方法

我们研究了栈勒伯格群体博弈(Stackelberg MFG)中的策略优化问题,该框架用于建模单一领导者与无限大同质追随者群体之间的层级战略互动。目标可表述为一个结构化的双层优化问题,领导者需学习最大化自身收益的策略,同时预判追随者的反应。现有方法通常依赖于领导者与追随者目标间严格的独立性假设,因嵌套循环结构导致采样效率低下,且缺乏有限时间收敛保证。为此,我们提出AC-SMFG,一种基于连续生成马尔可夫样本的单循环演员-评论家算法。该算法在领导者、代表性追随者及群体均值之间交替进行(半)梯度更新,实际实现简单。我们建立了该算法在有限时间和有限样本下的收敛性,证明其能收敛至栈勒伯格目标的平稳点。据我们所知,这是首个具有非渐近收敛保证的栈勒伯格群体博弈算法。关键假设为‘梯度对齐’条件,即领导者完整策略梯度可由其部分分量近似,放松了以往的独立性假设。在一系列经典经济学环境中的模拟结果表明,AC-SMFG在策略质量与收敛速度上均优于现有的多智能体与群体博弈学习基线。

原文摘要 · Abstract (English)

We study policy optimization in Stackelberg mean field games (MFGs), a hierarchical framework for modeling the strategic interaction between a single leader and an infinitely large population of homogeneous followers. The objective can be formulated as a structured bi-level optimization problem, in which the leader needs to learn a policy maximizing its reward, anticipating the response of the followers. Existing methods for solving these (and related) problems often rely on restrictive independence assumptions between the leader's and followers' objectives, use samples inefficiently due to nested-loop algorithm structure, and lack finite-time convergence guarantees. To address these limitations, we propose AC-SMFG, a single-loop actor-critic algorithm that operates on continuously generated Markovian samples. The algorithm alternates between (semi-)gradient updates for the leader, a representative follower, and the mean field, and is simple to implement in practice. We establish the finite-time and finite-sample convergence of the algorithm to a stationary point of the Stackelberg objective. To our knowledge, this is the first Stackelberg MFG algorithm with non-asymptotic convergence guarantees. Our key assumption is a "gradient alignment" condition, which requires that the full policy gradient of the leader can be approximated by a partial component of it, relaxing the existing leader-follower independence assumption. Simulation results in a range of well-established economics environments demonstrate that AC-SMFG outperforms existing multi-agent and MFG learning baselines in policy quality and convergence speed.

博弈学习群体博弈非渐近分析强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。