arXiv:2506.04251cs.AIcs.LG2025-06被引 3

用大模型提升多智能体协作,让机器人更会沟通和分工。

Language-Driven Coordination and Learning in Multi-Agent Simulation Environments

  • 引入大模型生成子目标、符号化通信与记忆回溯,增强协作能力。
  • 在足球、战争、星际等环境中,胜率和零样本泛化能力超越主流方法。
  • 适合研究智能体协作、人机协同或想用大模型做游戏训练的开发者。

本文提出 LLM-MARL 框架,将大语言模型(LLM)融入多智能体强化学习(MARL),以提升模拟游戏环境中的协作、通信与泛化能力。框架包含协调器、通信器和记忆模块,可动态生成子目标、实现符号化跨智能体通信,并支持情景记忆。训练结合 PPO 与语言条件损失及 LLM 查询门控机制。在 Google Research Football、MAgent Battle 与 StarCraft II 上评估,结果表明其在胜率、协作得分与零样本泛化方面均优于 MAPPO 与 QMIX。消融实验显示子目标生成与语言通信均显著贡献性能提升。定性分析发现角色分化与通信驱动战术等涌现行为。该工作推动了语言建模与策略学习融合,为交互式仿真中智能体设计提供了新路径,适用于训练、游戏与人机协作场景。

原文摘要 · Abstract (English)

This paper introduces LLM-MARL, a unified framework that incorporates large language models (LLMs) into multi-agent reinforcement learning (MARL) to enhance coordination, communication, and generalization in simulated game environments. The framework features three modular components of Coordinator, Communicator, and Memory, which dynamically generate subgoals, facilitate symbolic inter-agent messaging, and support episodic recall. Training combines PPO with a language-conditioned loss and LLM query gating. LLM-MARL is evaluated in Google Research Football, MAgent Battle, and StarCraft II. Results show consistent improvements over MAPPO and QMIX in win rate, coordination score, and zero-shot generalization. Ablation studies demonstrate that subgoal generation and language-based messaging each contribute significantly to performance gains. Qualitative analysis reveals emergent behaviors such as role specialization and communication-driven tactics. By bridging language modeling and policy learning, this work contributes to the design of intelligent, cooperative agents in interactive simulations. It offers a path forward for leveraging LLMs in multi-agent systems used for training, games, and human-AI collaboration.

多智能体大模型协作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。