梳理多智能体通信的五个核心问题,贯通强化学习到大模型的演进路径。
The Five Ws of Multi-Agent Communication: Who Talks to Whom, When, What, and Why -- A Survey from MARL to Emergent Language and LLMs
- 从手写协议到自学习语言,分三阶段解析通信机制演化。
- 揭示通信在动态环境中的关键作用:降低不确定性,提升协作效率。
- 适合研究多智能体系统、语言生成与协同控制的学者参考。
多智能体序列决策驱动众多现实系统,从自动驾驶到协作AI助手。在动态、部分可观测环境中,通信常是降低不确定性、实现协作的关键。本文通过‘谁对谁’、‘何时’、‘说什么’和‘为何有效’的五问框架,综述多智能体通信(MA-Comm)。该框架串联起原本分散的研究脉络。我们追溯通信方法在多智能体强化学习(MARL)、涌现语言(EL)和基于大语言模型(LLMs)三类范式中的演进:早期MARL采用人工设计或隐式协议,随后发展出端到端学习的通信以优化奖励与控制;虽有效但任务特定且难解释,推动了能通过交互生成结构化符号语言的涌现语言研究;然而其仍面临语义接地、泛化与可扩展性挑战,促使研究转向引入自然语言先验的大语言模型,用于更开放场景下的推理、规划与协作。本文提炼不同选择对通信设计的影响,揭示主要权衡与未解难题,总结实用设计模式与开放挑战,助力未来融合学习、语言与控制的可扩展、可解释多智能体协同系统。
原文摘要 · Abstract (English)
Multi-agent sequential decision-making powers many real-world systems, from autonomous vehicles and robotics to collaborative AI assistants. In dynamic, partially observable environments, communication is often what reduces uncertainty and makes collaboration possible. This survey reviews multi-agent communication (MA-Comm) through the Five Ws: who communicates with whom, what is communicated, when communication occurs, and why communication is beneficial. This framing offers a clean way to connect ideas across otherwise separate research threads. We trace how communication approaches have evolved across three major paradigms. In Multi-Agent Reinforcement Learning (MARL), early methods used hand-designed or implicit protocols, followed by end-to-end learned communication optimized for reward and control. While successful, these protocols are frequently task-specific and hard to interpret, motivating work on Emergent Language (EL), where agents can develop more structured or symbolic communication through interaction. EL methods, however, still struggle with grounding, generalization, and scalability, which has fueled recent interest in large language models (LLMs) that bring natural language priors for reasoning, planning, and collaboration in more open-ended settings. Across MARL, EL, and LLM-based systems, we highlight how different choices shape communication design, where the main trade-offs lie, and what remains unsolved. We distill practical design patterns and open challenges to support future hybrid systems that combine learning, language, and control for scalable and interpretable multi-agent collaboration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。