arXiv:2410.17898cs.LGcs.MA2024-10被引 2

离线强化学习解决大规模多智能体博弈均衡问题

Scalable Offline Reinforcement Learning for Mean Field Games

  • 基于离线数据与镜面下降法估计群体分布
  • 在人群探索任务中表现优于现有方法
  • 适合无法在线实验的真实多智能体场景

均值场强化学习为大规模交互智能体的策略优化提供了可扩展框架。现有方法通常依赖在线交互或系统动态模型,限制了在真实场景中的应用。本文提出离线蒙查乌森镜面下降(Off-MMD)算法,仅使用静态数据即可近似均值场博弈中的均衡策略。通过迭代镜面下降和重要性采样技术,该方法从离线数据中估计均值场分布,无需仿真或环境动态。同时结合离线强化学习技术,缓解Q值过估计问题,确保在数据覆盖有限时仍能稳健学习。算法在人群探索、导航等基准任务中表现优异,展示了在无法进行在线实验的真实多智能体系统中的适用性。实验验证了其对低质量数据的鲁棒性,并分析了超参数敏感性。

原文摘要 · Abstract (English)

Reinforcement learning algorithms for mean-field games offer a scalable framework for optimizing policies in large populations of interacting agents. Existing methods often depend on online interactions or access to system dynamics, limiting their practicality in real-world scenarios where such interactions are infeasible or difficult to model. In this paper, we present Offline Munchausen Mirror Descent (Off-MMD), a novel mean-field RL algorithm that approximates equilibrium policies in mean-field games using purely offline data. By leveraging iterative mirror descent and importance sampling techniques, Off-MMD estimates the mean-field distribution from static datasets without relying on simulation or environment dynamics. Additionally, we incorporate techniques from offline reinforcement learning to address common issues like Q-value overestimation, ensuring robust policy learning even with limited data coverage. Our algorithm scales to complex environments and demonstrates strong performance on benchmark tasks like crowd exploration or navigation, highlighting its applicability to real-world multi-agent systems where online experimentation is infeasible. We empirically demonstrate the robustness of Off-MMD to low-quality datasets and conduct experiments to investigate its sensitivity to hyperparameter choices.

多智能体离线学习均值场强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。