构建首个标准化的均场博弈学习评估套件,解决多智能体算法评测碎片化问题。
Bench-MFG: A Benchmark Suite for Learning in Stationary Mean Field Games
- 提出分层问题分类体系,覆盖无交互、势能型到动态耦合等典型博弈场景。
- 设计可随机生成的MF-Garnets环境,支持大规模统计测试与算法鲁棒性验证。
- 评测多种学习算法并给出统一实验标准,适合多智能体强化学习研究者参考。
均场博弈(MFG)与强化学习(RL)的融合催生了大量求解大规模多智能体系统的方法,但当前缺乏统一的评估协议,导致研究者依赖自定义、孤立且简化的环境,难以评估方法的鲁棒性、泛化能力及失效模式。为此,我们提出了一个面向离散时间、离散状态、静态设定的均场博弈综合基准套件(Bench-MFG)。我们建立了一套问题分类体系,涵盖无交互、单调博弈、势能博弈和动态耦合博弈等类型,并为每类提供典型环境实例。此外,提出MF-Garnets方法以生成随机均场博弈实例,支持严谨的统计测试。我们在这些环境中对多种学习算法进行了基准测试,包括一种新的黑箱式剥削最小化方法(MF-PSO)。基于广泛的实证结果,我们提出未来实验比较的标准化建议。代码已公开于 https://github.com/lorenzomagnino/Bench-MFG。
原文摘要 · Abstract (English)
The intersection of Mean Field Games (MFGs) and Reinforcement Learning (RL) has fostered a growing family of algorithms designed to solve large-scale multi-agent systems. However, the field currently lacks a standardized evaluation protocol, forcing researchers to rely on bespoke, isolated, and often simplistic environments. This fragmentation makes it difficult to assess the robustness, generalization, and failure modes of emerging methods. To address this gap, we propose a comprehensive benchmark suite for MFGs (Bench-MFG), focusing on the discrete-time, discrete-space, stationary setting for the sake of clarity. We introduce a taxonomy of problem classes, ranging from no-interaction and monotone games to potential and dynamics-coupled games, and provide prototypical environments for each. Furthermore, we propose MF-Garnets, a method for generating random MFG instances to facilitate rigorous statistical testing. We benchmark a variety of learning algorithms across these environments, including a novel black-box approach (MF-PSO) for exploitability minimization. Based on our extensive empirical results, we propose guidelines to standardize future experimental comparisons. Code available at \href{https://github.com/lorenzomagnino/Bench-MFG}{https://github.com/lorenzomagnino/Bench-MFG}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。