提出M2I2框架,让多智能体更高效地理解与利用通信信息。
M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference
- 通过掩码状态建模和意图推理增强信息吸收能力。
- 在多种复杂场景中超越现有方法,提升通信效率与泛化性能。
- 适合研究多智能体协作、通信优化与决策协同的学者。
通信在协调多智能体行为中至关重要。然而,现有方法多关注信息内容、发送时机与接收方,忽视了信息整合能力,影响智能体对复杂不确定交互的理解与响应,从而降低整体通信效率。为此,我们提出M2I2框架,旨在提升智能体有效吸收与利用接收信息的能力。M2I2赋予智能体先进的掩码状态建模与联合动作预测能力,增强对环境不确定性的感知,并促进对队友意图的预判。该机制确保智能体获得全面且相关的信息,推动更明智、协同的行为。此外,我们设计了一种基于维度理性的网络,通过元学习训练,识别各维度信息的重要性,评估其对决策与辅助任务的贡献。进而采用重要性驱动的启发式策略进行选择性信息掩码与共享,优化掩码状态建模效率与信息共享逻辑。我们在多个多智能体任务上评估M2I2,结果表明其在复杂场景中显著优于现有先进方法,在性能、效率与泛化能力方面均具优势。
原文摘要 · Abstract (English)
Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。