根据需求定制通信,让智能体更高效协作。
DCMAC: Demand-aware Customized Multi-Agent Communication via Upper Bound Training
- 用上界训练获得理想策略,指导通信决策。
- 通信效率提升,约束与非约束场景均优于基线。
- 适合资源受限的多智能体协同任务。
高效的通信能提升多智能体强化学习的整体性能。传统方法通过全量信息共享,导致通信开销过大。现有工作尝试基于局部信息推断全局状态,但忽略了预测不确定性带来的训练困难。为此,本文提出需求感知的定制化多智能体通信协议DCMAC,通过上界训练获取理想策略。利用需求解析模块,智能体可评估向队友发送本地信息的收益,并通过交叉注意力机制计算需求与本地观测的相关性,生成定制消息。此外,该方法可适应不同通信资源,借助联合观测训练的理想策略加速训练进程。实验表明,DCMAC在无约束和通信受限场景下均显著优于基线算法。
原文摘要 · Abstract (English)
Efficient communication can enhance the overall performance of collaborative multi-agent reinforcement learning. A common approach is to share observations through full communication, leading to significant communication overhead. Existing work attempts to perceive the global state by conducting teammate model based on local information. However, they ignore that the uncertainty generated by prediction may lead to difficult training. To address this problem, we propose a Demand-aware Customized Multi-Agent Communication (DCMAC) protocol, which use an upper bound training to obtain the ideal policy. By utilizing the demand parsing module, agent can interpret the gain of sending local message on teammate, and generate customized messages via compute the correlation between demands and local observation using cross-attention mechanism. Moreover, our method can adapt to the communication resources of agents and accelerate the training progress by appropriating the ideal policy which is trained with joint observation. Experimental results reveal that DCMAC significantly outperforms the baseline algorithms in both unconstrained and communication constrained scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。