arXiv:2604.03785cs.AIcs.MA2026-04

提出通信增益与延迟成本度量,提升延迟通信下的多智能体协作性能。

Communication Gain and Delay Cost Under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

  • 用通信增益与延迟成本量化消息价值,指导何时发送与处理消息。
  • 在多个延迟水平下,任务成功率提升12%-28%,泛化能力显著增强。
  • 适合研究多智能体系统中通信优化的开发者与研究人员参考。

在部分可观测的协作多智能体强化学习中,通信对协调至关重要,但跨时间步延迟导致信息传递滞后,造成时间错配和信息过时。本文将该场景形式化为延迟通信的部分可观测马尔可夫博弈(DeComm-POMG),并将消息的影响分解为通信增益与延迟成本,提出CGDC度量。进一步建立了值损失上界,表明延迟带来的性能退化由及时与延迟消息诱导的动作分布间信息差距的折扣累积所限制。基于CGDC,提出CDCMA框架:仅当预测通信增益大于延迟成本时请求消息;利用未来观测预测减少消费时的错配;通过CGDC引导的注意力融合延迟消息。在无队友视野的协作导航与猎物追捕任务,以及SMAC多地图、多延迟水平下的实验均显示性能、鲁棒性与泛化能力持续提升,消融实验验证了各组件有效性。

原文摘要 · Abstract (English)

Communication is essential for coordination in \emph{cooperative} multi-agent reinforcement learning under partial observability, yet \emph{cross-timestep} delays cause messages to arrive multiple timesteps after generation, inducing temporal misalignment and making information stale when consumed. We formalize this setting as a delayed-communication partially observable Markov game (DeComm-POMG) and decompose a message's effect into \emph{communication gain} and \emph{delay cost}, yielding the Communication Gain and Delay Cost (CGDC) metric. We further establish a value-loss bound showing that the degradation induced by delayed messages is upper-bounded by a discounted accumulation of an information gap between the action distributions induced by timely versus delayed messages. Guided by CGDC, we propose \textbf{CDCMA}, an actor--critic framework that requests messages only when predicted CGDC is positive, predicts future observations to reduce misalignment at consumption, and fuses delayed messages via CGDC-guided attention. Experiments on no-teammate-vision variants of Cooperative Navigation and Predator Prey, and on SMAC maps across multiple delay levels show consistent improvements in performance, robustness, and generalization, with ablations validating each component.

多智能体强化学习通信优化延迟建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。