让分散智能体在无奖励依赖下有效通信,提升协作效率。
Reward-Independent Messaging for Decentralized Multi-Agent Reinforcement Learning
- 用预测编码机制建模消息,不依赖合作或共享参数
- 两种算法在非合作任务中均超越传统消息作为动作的方法
- 适合复杂分散系统中的协调问题,如多机器人、网络控制
在多智能体强化学习(MARL)中,有效通信能显著提升部分可观测环境下的性能。我们提出MARL-CPC框架,使完全去中心化、独立的智能体无需参数共享即可通信。该框架基于涌现通信研究中的集体预测编码(CPC)构建消息学习模型。与将消息视为动作空间一部分并假设合作的传统方法不同,MARL-CPC将消息与状态推断关联,支持在非合作、奖励无关场景下的通信。我们提出两种算法——Bandit-CPC和IPPO-CPC,并在非合作MARL任务中进行评估。基准测试显示,两者均优于标准消息作为动作的方法,即使消息对发送者无直接收益也能实现有效通信。结果表明,MARL-CPC在复杂去中心化环境中具备实现协调的潜力。
原文摘要 · Abstract (English)
In multi-agent reinforcement learning (MARL), effective communication improves agent performance, particularly under partial observability. We propose MARL-CPC, a framework that enables communication among fully decentralized, independent agents without parameter sharing. MARL-CPC incorporates a message learning model based on collective predictive coding (CPC) from emergent communication research. Unlike conventional methods that treat messages as part of the action space and assume cooperation, MARL-CPC links messages to state inference, supporting communication in non-cooperative, reward-independent settings. We introduce two algorithms -Bandit-CPC and IPPO-CPC- and evaluate them in non-cooperative MARL tasks. Benchmarks show that both outperform standard message-as-action approaches, establishing effective communication even when messages offer no direct benefit to the sender. These results highlight MARL-CPC's potential for enabling coordination in complex, decentralized environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。