arXiv:2410.07976stat.MLcs.LG2024-10被引 3

用变分不等式解决多智能体强化学习中的旋转优化问题

Addressing Rotational Learning Dynamics in Multi-Agent Reinforcement Learning

  • 将多智能体强化学习重构成变分不等式框架,统一处理竞争目标的旋转动态
  • 在石头剪刀布和匹配硬币游戏中收敛到均衡策略,提升团队协作能力
  • 可嵌入现有算法,适合研究复杂协作与对抗场景的学者使用

多智能体强化学习(MARL)通过智能体间的合作与竞争解决复杂问题,应用广泛。尽管取得成功,但存在可复现性危机。我们发现,部分原因在于竞争目标引发的旋转优化动态,需超越传统优化方法。本文将MARL重构为变分不等式(VIs)框架,提出一种通用方法,集成基于梯度的VI优化技术以应对旋转动态。实验表明,在零和博弈如石头剪刀布和匹配硬币中,VI方法实现更好均衡收敛;在多智能体粒子环境Predator-prey中,显著提升团队协作表现。结果凸显先进优化技术对MARL的变革潜力。

原文摘要 · Abstract (English)

Multi-agent reinforcement learning (MARL) has emerged as a powerful paradigm for solving complex problems through agents' cooperation and competition, finding widespread applications across domains. Despite its success, MARL faces a reproducibility crisis. We show that, in part, this issue is related to the rotational optimization dynamics arising from competing agents' objectives, and require methods beyond standard optimization algorithms. We reframe MARL approaches using Variational Inequalities (VIs), offering a unified framework to address such issues. Leveraging optimization techniques designed for VIs, we propose a general approach for integrating gradient-based VI methods capable of handling rotational dynamics into existing MARL algorithms. Empirical results demonstrate significant performance improvements across benchmarks. In zero-sum games, Rock--paper--scissors and Matching pennies, VI methods achieve better convergence to equilibrium strategies, and in the Multi-Agent Particle Environment: Predator-prey, they also enhance team coordination. These results underscore the transformative potential of advanced optimization techniques in MARL.

多智能体强化学习优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。