arXiv:2409.20067cs.LGcs.GT2024-09被引 6

提出新型鲁棒多智能体强化学习框架,突破样本效率瓶颈。

Breaking the Curse of Multiagency in Robust Multi-Agent Reinforcement Learning

  • 基于行为经济学构建动态不确定性集,融合环境与协同行为
  • 首个实现多项式样本复杂度的鲁棒多智能体算法
  • 适用于存在对抗扰动的现实场景,适合系统设计者参考

标准多智能体强化学习算法易受仿真到现实的差距影响。为此,分布鲁棒马尔可夫博弈(RMG)通过优化在预设不确定性集内动态变化时的最坏情况性能,提升鲁棒性。然而,RMG仍缺乏合理问题建模和高效算法。两大核心挑战是不确定性集的合理定义,以及是否能克服多智能体的“诅咒”——即样本复杂度随智能体数量呈指数增长。本文提出一类受行为经济学启发的自然RMG类,其中每个智能体的不确定性集由环境及其他智能体的整体行为共同决定。我们首先证明该类RMG的适定性,建立了鲁棒纳什均衡和粗相关均衡(CCE)的存在性。在生成模型假设下,提出一种样本高效的CCE学习算法,其样本复杂度关于所有相关参数呈多项式增长。据我们所知,这是首个在任意不确定性集设定下打破多智能体诅咒的算法。

原文摘要 · Abstract (English)

Standard multi-agent reinforcement learning (MARL) algorithms are vulnerable to sim-to-real gaps. To address this, distributionally robust Markov games (RMGs) have been proposed to enhance robustness in MARL by optimizing the worst-case performance when game dynamics shift within a prescribed uncertainty set. RMGs remains under-explored, from reasonable problem formulation to the development of sample-efficient algorithms. Two notorious and open challenges are the formulation of the uncertainty set and whether the corresponding RMGs can overcome the curse of multiagency, where the sample complexity scales exponentially with the number of agents. In this work, we propose a natural class of RMGs inspired by behavioral economics, where each agent's uncertainty set is shaped by both the environment and the integrated behavior of other agents. We first establish the well-posedness of this class of RMGs by proving the existence of game-theoretic solutions such as robust Nash equilibria and coarse correlated equilibria (CCE). Assuming access to a generative model, we then introduce a sample-efficient algorithm for learning the CCE whose sample complexity scales polynomially with all relevant parameters. To the best of our knowledge, this is the first algorithm to break the curse of multiagency for RMGs, regardless of the uncertainty set formulation.

多智能体鲁棒学习强化学习博弈论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。