arXiv:2501.12199cs.LGcs.GT2025-01被引 3

提出基于创新动态的经验回放算法,提升多智能体强化学习稳定性。

Experience-replay Innovative Dynamics

  • 用可调超参数的修正协议替代传统策略,模拟创新动态轨迹。
  • 在非稳定博弈中实现周期性行为,逼近纳什均衡。
  • 为多智能体强化学习提供超越复制者动态的新理论框架,适合研究博弈收敛的学者。

尽管多智能体强化学习(MARL)取得突破性进展,但仍面临不稳定和非平稳性问题。演化博弈论中的复制者动态能保证稳定博弈下的收敛性,但在其他场景下表现相反,亟需替代方案。相比之下,如布朗-冯诺依曼-纳什(BNN)或史密斯等创新动态会产生周期性轨迹,具备逼近纳什均衡的潜力,但尚未有基于此类动态的MARL算法。为此,本文提出一种新型基于经验回放的MARL算法,将修正协议作为可调超参数引入。通过合理调整协议,算法行为可复现这些动态的轨迹。更重要的是,该工作构建了一个可扩展的理论框架,使MARL算法的收敛性保证超越复制者动态。最后,实验验证了理论分析的有效性。

原文摘要 · Abstract (English)

Despite its groundbreaking success, multi-agent reinforcement learning (MARL) still suffers from instability and nonstationarity. Replicator dynamics, the most well-known model from evolutionary game theory (EGT), provide a theoretical framework for the convergence of the trajectories to Nash equilibria and, as a result, have been used to ensure formal guarantees for MARL algorithms in stable game settings. However, they exhibit the opposite behavior in other settings, which poses the problem of finding alternatives to ensure convergence. In contrast, innovative dynamics, such as the Brown-von Neumann-Nash (BNN) or Smith, result in periodic trajectories with the potential to approximate Nash equilibria. Yet, no MARL algorithms based on these dynamics have been proposed. In response to this challenge, we develop a novel experience replay-based MARL algorithm that incorporates revision protocols as tunable hyperparameters. We demonstrate, by appropriately adjusting the revision protocols, that the behavior of our algorithm mirrors the trajectories resulting from these dynamics. Importantly, our contribution provides a framework capable of extending the theoretical guarantees of MARL algorithms beyond replicator dynamics. Finally, we corroborate our theoretical findings with empirical results.

多智能体强化学习博弈论动态系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。