提出分层消息传递策略,提升多智能体协作与规划能力。
Hierarchical Message-Passing Policies for Multi-Agent Reinforcement Learning
- 采用封建式分层强化学习框架,通过图结构实现层级协调
- 下层智能体接收上层目标并同级通信,显著提升决策效率
- 新奖励分配机制使下层策略最大化上层优势函数,效果优于现有方法
去中心化多智能体强化学习虽具可扩展性,但面临局部观测和非平稳性挑战。本文提出一种新型多智能体分层消息传递策略,基于封建式强化学习框架,利用分层图结构实现智能体间的规划与协调。下层智能体接收上层目标,并与同级邻居交换消息。为训练分层策略,设计了一种新奖励分配机制:使下层策略最大化与上层相关的优势函数。在多个基准测试中,该方法表现优于当前最优水平。
原文摘要 · Abstract (English)
Decentralized Multi-Agent Reinforcement Learning (MARL) methods allow for learning scalable multi-agent policies, but suffer from partial observability and induced non-stationarity. These challenges can be addressed by introducing mechanisms that facilitate coordination and high-level planning. Specifically, coordination and temporal abstraction can be achieved through communication (e.g., message passing) and Hierarchical Reinforcement Learning (HRL) approaches to decision-making. However, optimization issues limit the applicability of hierarchical policies to multi-agent systems. As such, the combination of these approaches has not been fully explored. To fill this void, we propose a novel and effective methodology for learning multi-agent hierarchies of message-passing policies. We adopt the feudal HRL framework and rely on a hierarchical graph structure for planning and coordination among agents. Agents at lower levels in the hierarchy receive goals from the upper levels and exchange messages with neighboring agents at the same level. To learn hierarchical multi-agent policies, we design a novel reward-assignment method based on training the lower-level policies to maximize the advantage function associated with the upper levels. Results on relevant benchmarks show that our method performs favorably compared to the state of the art.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。