让智能体学会合作规律,提升多智能体协作效率。
LLawCo: Learning Laws of Cooperation for Modeling Embodied Multi-Agent Behavior

- 通过反思失败经验提取合作行为规律,显式融入推理链
- 在两个基准上平均提升6.8%任务成功率
- 适合研究多智能体协作与语言模型对齐的学者
在去中心化和部分可观测环境中运行的具身智能体近年来受到广泛关注。然而,现有基于大语言模型(LLM)的智能体常表现出与伙伴不协调或与环境状态不一致的行为,导致协作效率低下、任务成功率差。为此,我们提出一种新框架——学习合作规律(LLawCo),使具身智能体能够自主地与伙伴及任务目标对齐。该框架让智能体反思过往失败,提取行为偏差模式,并归纳出高层行为规律,如“必要时才说话”“等待伙伴”。这些规律通过监督微调显式嵌入智能体的思维链中,使其推理更契合任务需求与他人行为。为评估该方法,我们构建了基于PARTNR环境的大规模多智能体对话与协作规划基准PARTNR-Dialog。在已有任务及新基准上的实验表明,合作效率与任务成功率显著提升。在四个骨干LLM上,本方法在PARTNR-Dialog基准上平均成功率提升4.5%,在TDW-MAT基准上提升6.8%,优于当前最先进的开源通信智能体框架。
原文摘要 · Abstract (English)
Embodied agents operating in decentralized and partially observable environments have attracted growing attention in recent years. However, existing large language model (LLM)-based agents often exhibit behaviors that are misaligned with their partners or inconsistent with the environment state, leading to inefficient cooperation and poor task success. To address this challenge, we propose a novel framework, Learning Laws of Cooperation (LLawCo), that enables embodied agents to autonomously align with both their partners and task objectives. Our framework allows agents to reflect on past failures to extract misaligned behavioral patterns, which are used to derive high-level behavioral laws, such as "Talk when necessary" and "Wait for partner." These laws are explicitly incorporated into the agents' chains of thought via supervised fine-tuning, aligning their reasoning with task requirements and the behavior of other agents. To evaluate our approach, we introduce PARTNR-Dialog, a large-scale multi-agent communicative and cooperative planning benchmark built on the PARTNR environment. Experiments on existing tasks and our new benchmark demonstrate significant improvements in cooperative efficiency and task success rates. Across four backbone LLMs, our method achieves average success rate improvements of 4.5% on the PARTNR-Dialog benchmark and 6.8% on the TDW-MAT benchmark over state-of-the-art open-source communicative agent frameworks. See the LLawCo project page for details: https://www.merl.com/research/highlights/LLawCo
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。