让多智能体学会在真实无线信道中高效通信并协作
Wireless Communication Enhanced Value Decomposition for Multi-Agent Reinforcement Learning

- 用图神经网络建模真实无线信道下的通信关系,指导价值分解
- 在捕食者-猎物和伐木工任务中,收敛更快且性能更高
- 适合研究无线环境中的多智能体协作与通信机制
多智能体强化学习中的协作依赖智能体间通信,但现有方法常假设理想通信信道,且价值分解忽略信息传递的谁与谁关系。本文提出CLOVER框架,其集中式价值混合器基于实际无线信道下的通信图进行条件化。该通信图引入关系归纳偏置,依据实际通信结构约束个体效用的混合方式。混合器采用节点权重由置换等变超网络生成的图神经网络:沿通信边的多跳传播重构信用分配,不同拓扑结构引发不同混合策略。理论证明该混合器具有置换不变性、单调性(保持IGM条件),且表达能力强于QMIX类方法。为应对真实信道,我们构建扩展的MDP以分离随机信道效应与智能体计算图,并使用随机接收域编码器处理可变大小的消息集合,支持端到端可微训练。在p-CSMA无线信道下的捕食者-猎物和伐木工基准测试中,CLOVER在收敛速度与最终性能上均优于VDN、QMIX、TarMAC+VDN和TarMAC+QMIX。行为分析表明智能体学会自适应的发送与监听策略,消融实验确认通信图归纳偏置是性能提升的关键。
原文摘要 · Abstract (English)
Cooperation in multi-agent reinforcement learning (MARL) benefits from inter-agent communication, yet most approaches assume idealized channels and existing value decomposition methods ignore who successfully shared information with whom. We propose CLOVER, a cooperative MARL framework whose centralized value mixer is conditioned on the communication graph realized under a realistic wireless channel. This graph introduces a relational inductive bias into value decomposition, constraining how individual utilities are mixed based on the realized communication structure. The mixer is a GNN with node-specific weights generated by a Permutation-Equivariant Hypernetwork: multi-hop propagation along communication edges reshapes credit assignment so that different topologies induce different mixing. We prove this mixer is permutation invariant, monotonic (preserving the IGM condition), and strictly more expressive than QMIX-style mixers. To handle realistic channels, we formulate an augmented MDP isolating stochastic channel effects from the agent computation graph, and employ a stochastic receptive field encoder for variable-size message sets, enabling end-to-end differentiable training. On Predator-Prey and Lumberjacks benchmarks under p-CSMA wireless channels, CLOVER consistently improves convergence speed and final performance over VDN, QMIX, TarMAC+VDN, and TarMAC+QMIX. Behavioral analysis confirms agents learn adaptive signaling and listening strategies, and ablations isolate the communication-graph inductive bias as the key source of improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。