提出新算法求解带公共噪声的群体博弈均衡,无需历史数据。
Population-aware Online Mirror Descent for Mean-Field Games with Common Noise by Deep Reinforcement Learning
- 结合强化学习与在线镜像下降,直接学习依赖群体分布的策略。
- 在7个典型场景中收敛速度优于现有最优方法,尤其在公共噪声下表现更优。
- 适合研究大规模多智能体系统,尤其适用于初始分布未知的情况。
平均场博弈(MFGs)为研究大规模多智能体系统提供了强大框架。然而,在初始分布未知或群体受公共噪声影响时,学习纳什均衡仍是难题。本文提出一种高效的深度强化学习(DRL)算法,旨在实现不依赖平均化或历史采样的群体相关纳什均衡,灵感来自Munchausen RL和在线镜像下降。所得策略可适应不同初始分布及各类公共噪声来源。在七个典型例子上的数值实验表明,该算法在收敛性上优于当前最先进的方法,特别是针对群体相关策略的虚构博弈DRL版本。在公共噪声环境中的表现凸显了本方法的鲁棒性与适应性。
原文摘要 · Abstract (English)
Mean Field Games (MFGs) offer a powerful framework for studying large-scale multi-agent systems. Yet, learning Nash equilibria in MFGs remains a challenging problem, particularly when the initial distribution is unknown or when the population is subject to common noise. In this paper, we introduce an efficient deep reinforcement learning (DRL) algorithm designed to achieve population-dependent Nash equilibria without relying on averaging or historical sampling, inspired by Munchausen RL and Online Mirror Descent. The resulting policy is adaptable to various initial distributions and sources of common noise. Through numerical experiments on seven canonical examples, we demonstrate that our algorithm exhibits superior convergence properties compared to state-of-the-art algorithms, particularly a DRL version of Fictitious Play for population-dependent policies. The performance in the presence of common noise underscores the robustness and adaptability of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。