在存在全局噪声的群体博弈中,学习能感知群体状态的策略,才能准确模仿最优行为。
Population-Aware Imitation Learning in Mean-field Games with Common Noise

- 设计两类学习目标:求解纳什均衡与逼近专家群体表现
- 证明最小化模仿代理可控制策略的可被利用性与性能差距
- 提出结合拟合策略与深度学习的数值框架,适用于真实场景
均值场博弈(MFGs)为大规模交互智能体的集体行为建模提供了强大框架。本文研究在存在全局噪声的MFG中进行模仿学习的问题,此时群体分布随时间随机演化,迫使智能体采用感知群体状态的策略以应对整体冲击。我们提出两种学习目标:恢复纳什均衡和最大化对专家群体的表现。研究了两种模仿代理:行为克隆(BC)和对抗性分歧(ADV)。我们建立了有限样本误差界,表明最小化这些代理能有效控制策略的可被利用性及其相对于专家的性能差距。此外,我们提出一种基于广义虚构演化的数值框架,结合深度学习来计算专家级的群体感知策略。在三个环境上的实验表明,标准的无群体感知策略无法捕捉均衡动态。结果强调,在存在全局噪声时,学习群体感知策略对于避免被随机性误导至关重要。
原文摘要 · Abstract (English)
Mean Field Games (MFGs) provide a powerful framework for modeling the collective behavior of large populations of interacting agents. In this paper, we address the problem of Imitation Learning (IL) in MFGs subject to common noise, where the population distribution evolves stochastically. This stochasticity compels agents to adopt population-aware policies to respond to aggregate shocks. We formulate two distinct learning objectives: recovering a Nash equilibrium and maximizing performance against an expert population. We investigate two imitation proxies: Behavioral Cloning (BC) and Adversarial (ADV) divergence. We then establish finite-sample error bounds showing that minimizing these proxies effectively controls both the policy's exploitability and its performance gap relative to the expert. Furthermore, we propose a numerical framework using generalized Fictitious Play and Deep Learning to compute expert population-aware policies. Through experiments on three environments we demonstrate that standard population-unaware policies fail to capture the equilibrium dynamics. Our results highlight that learning population-aware policies is crucial to avoid being misled by the randomness inherent in common noise.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。