让大模型个体学习、团队协作进化,提升智能体在真实环境中的合作能力。
Learn as Individuals, Evolve as a Team: Multi-agent LLMs Adaptation in Embodied Environments
- 个体学习环境知识,团队共建共享协作经验
- 在两个基准上优于现有方法,显著提升协作规划效果
- 适合需要动态协作的多智能体系统研究者
大语言模型(LLMs)具备丰富的知识和强大的推理能力,是复杂多智能体规划任务的理想工具。然而,现有基于LLM的规划算法在多智能体具身环境中仍受限于适应能力弱的问题。为此,本文提出一种名为「个体学习、团队进化(LIET)」的框架,使LLM智能体在训练前和测试期间均能持续学习与演化。在个体层面,智能体通过探索数据集学习局部效用函数,以理解具身环境,并在测试中用于辅助决策;在团队层面,智能体协同迭代更新共享的协作知识列表,指导更有效的沟通。结合个体学习与团队演化,实现全面而灵活的适应性。在Communicative Watch-And-Help与ThreeD-World Multi-Agent Transport两个基准上的实验表明,无论是使用LLaMA还是GPT-4o,LIET均显著超越现有基线,展现出强大的协作规划能力。
原文摘要 · Abstract (English)
Large language models (LLMs) possess extensive knowledge bases and strong reasoning capabilities, making them promising tools for complex, multi-agent planning in embodied environments. However, despite LLMs' advanced abilities and the sophisticated modular design of agentic methods, existing LLM-based planning algorithms remain limited by weak adaptation capabilities to multi-agent embodied scenarios. We address this limitation by introducing a framework that enables LLM agents to learn and evolve both before and during test time, equipping them with environment-relevant knowledge for better planning and enhanced communication for improved cooperation. Inspired by centralized training with decentralized execution in multi-agent reinforcement learning, we propose a \textit{Learn as Individuals, Evolve as a Team (LIET)} paradigm for multi-agent LLMs adaptation. At the individual level, LLM agents learn a local utility function from exploratory datasets to better comprehend the embodied environment, which is then queried during test time to support informed decision-making. At the team level, LLM agents collaboratively and iteratively maintain and update a shared cooperation knowledge list based on new experiences, using it to guide more effective communication. By combining individual learning with team evolution, LIET enables comprehensive and flexible adaptation for LLM agents. Our experiments on Communicative Watch-And-Help and ThreeD-World Multi-Agent Transport benchmarks demonstrate that LIET, instantiated with both LLaMA and GPT-4o, outperforms existing baselines and exhibits strong cooperative planning abilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。