让多个智能体按优先级通信,提升决策质量与协作效率。
Multi-Agent Decision-Focused Learning via Value-Aware Sequential Communication

- 智能体按顺序发送消息,后发者根据前序者信息调整策略。
- 在医疗和星际争霸任务中,奖励提升4~6倍,胜率提高13%以上。
- 适合需要高效协作的多智能体系统,如机器人集群、智能调度。
在部分可观测环境下,多智能体协作需共享互补的私有信息。现有方法优化消息以达成中间目标(如重构准确率或互信息),而非直接提升决策质量。本文提出 SeqComm-DFL,将序列化通信与决策导向学习统一。其核心为价值感知的消息生成与序列化斯塔克尔伯格条件:消息按优先级生成,后继智能体基于前序者信息进行条件化决策,其引导潜力由利他性排序决定。我们扩展最优模型设计至通信增强型世界模型,并采用 QMIX 因子分解,实现通过隐式微分的端到端高效训练。理论证明通信价值随协调差距增长,且双层优化收敛速度为 $/mathcal{O}(1/ ext{sqrt}{T})$,其中 $T$ 为训练迭代次数。在协作医疗与星际争霸多智能体挑战(SMAC)基准上,SeqComm-DFL 实现累计奖励提升4~6倍,胜率提升超13%,实现了信息不对称下不可达的协作策略。
原文摘要 · Abstract (English)
Multi-agent coordination under partial observability requires agents to share complementary private information. While recent methods optimize messages for intermediate objectives (e.g., reconstruction accuracy or mutual information), rather than decision quality, we introduce \textbf{SeqComm-DFL}, unifying the sequential communication with decision-focused learning for task performance. Our approach features \emph{value-aware message generation with sequential Stackelberg conditioning}: messages maximize receiver decision quality and are generated in priority order, with agents conditioning on their predecessors. The \emph{guidance potential} determined by their prosocial ordering. We extend Optimal Model Design to communication-augmented world models with QMIX factorization, enabling efficient end-to-end training via implicit differentiation. We prove information-theoretic bounds showing that communication value scales with coordination gaps and establish $\mathcal{O}(1/\sqrt{T})$ convergence for the bilevel optimization, where $T$ denotes the number of training iterations. On collaborative healthcare and StarCraft Multi-Agent Challenge (SMAC) benchmarks, SeqComm-DFL achieves four to six times higher cumulative rewards and over 13\% win rate improvements, enabling coordination strategies inaccessible under information asymmetry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。