提出新框架提升无通信多智能体协作能力
Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial Exploration
- 通过信念建模推断不可观测状态,过滤冗余信息
- 结合信念增强策略判别力,对抗性探索发现高价值状态
- 在多个基准测试中超越现有最优算法
在无通信的分布式部分可观测环境中实现协作,对多智能体深度强化学习(MARL)构成重大挑战。本文聚焦于从个体智能体观测中推断状态表示,并利用这些表示增强智能体的探索与协作策略。为此,提出一种新型协作MARL状态建模框架,使智能体基于自身策略优化目标,推断出有意义的非可观测状态信念,同时过滤冗余和低信息量的联合状态信息。在此框架基础上,提出SMPE算法:智能体通过将信念显式融入策略网络,提升自身在部分可观测条件下的判别能力;并通过采用对抗性探索策略,主动发现新颖且高价值状态,同时提升其他智能体的判别能力。实验表明,SMPE在MPE、LBF和RWARE基准中的复杂全协作任务上,显著优于当前最先进的MARL算法。
原文摘要 · Abstract (English)
Learning to cooperate in distributed partially observable environments with no communication abilities poses significant challenges for multi-agent deep reinforcement learning (MARL). This paper addresses key concerns in this domain, focusing on inferring state representations from individual agent observations and leveraging these representations to enhance agents' exploration and collaborative task execution policies. To this end, we propose a novel state modelling framework for cooperative MARL, where agents infer meaningful belief representations of the non-observable state, with respect to optimizing their own policies, while filtering redundant and less informative joint state information. Building upon this framework, we propose the MARL SMPE algorithm. In SMPE, agents enhance their own policy's discriminative abilities under partial observability, explicitly by incorporating their beliefs into the policy network, and implicitly by adopting an adversarial type of exploration policies which encourages agents to discover novel, high-value states while improving the discriminative abilities of others. Experimentally, we show that SMPE outperforms state-of-the-art MARL algorithms in complex fully cooperative tasks from the MPE, LBF, and RWARE benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。