提出新表征框架,让多智能体在不传消息情况下也能默契配合。
Unifying Agent Interaction and World Information for Multi-agent Coordination
- 通过建模通信协议,联合学习智能体间关系与任务世界信息
- 在4个复杂基准上表现优异,支持去中心化执行且隐式协调
- 可兼容现有MARL算法,适合需要高效协同的系统设计
本文提出一种新型表征学习框架——交互-世界潜在(IWoL),以促进多智能体强化学习(MARL)中的团队协作。由于多智能体交互带来的复杂动态以及局部观测导致的信息不完整,构建有效的团队协作表征极具挑战。核心洞察在于,通过直接建模通信协议,构建一个可学习的表征空间,联合捕捉智能体间关系与任务相关的世界信息。该表征支持完全去中心化执行,并实现隐式协调,避免了显式消息传递的缺点,如决策延迟、易受恶意攻击及对带宽敏感。实际应用中,该表征既可作为每个智能体的隐式潜在变量,也可作为显式通信消息。在四个具有挑战性的MARL基准上评估两种变体,结果表明IWoL为团队协作提供了一种简单而强大的方案。此外,我们证明该表征可与现有MARL算法结合,进一步提升其性能。
原文摘要 · Abstract (English)
This work presents a novel representation learning framework, *interaction-world* latent (IWoL), to facilitate *team coordination* in multi-agent reinforcement learning (MARL). Building effective representation for team coordination is a challenging problem, due to the intricate dynamics emerging from multi-agent interaction and incomplete information induced by local observations. Our key insight is to construct a learnable representation space that jointly captures inter-agent relations and task-specific world information by directly modeling communication protocols. This representation enables fully decentralized execution with implicit coordination while avoiding the drawbacks of explicit message passing, for example, slower decision-making, vulnerability to malicious attackers, and sensitivity to bandwidth limitations. In practice, our representation can be used not only as an implicit latent for each agent, but also as an explicit message for communication. Across four challenging MARL benchmarks, we evaluate both variants and show that IWoL provides a simple yet powerful key for team coordination. Moreover, we demonstrate that our representation can be combined with existing MARL algorithms to further enhance their performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。