用游戏测试大模型协作能力,发现它们懂任务却难配合。
Collab-Overcooked: Benchmarking and Evaluating Large Language Models as Collaborative Agents
- 在烹饪游戏环境中让大模型通过语言协作完成任务
- 13个模型中多数能理解目标但缺乏持续协同与适应能力
- 提供可复现的评测框架,适合研究多智能体系统者使用
基于大语言模型的智能体系统在现实应用中取得显著进展。本文提出一个新的基于大语言模型的多智能体系统(LLM-MAS)基准——Collab-Overcooked,构建于流行的Overcooked-AI游戏之上,包含更贴近真实场景且更具挑战性的交互任务。该基准在两方面实现创新:一是提供支持多样化任务与目标的多智能体框架,并通过自然语言通信促进协作;二是引入一系列过程导向的评估指标,用于细致评估不同大模型智能体的协作能力,这一维度在以往研究中常被忽略。我们对13个主流大语言模型进行了广泛实验,结果显示,尽管模型在任务目标理解上表现良好,但在主动协作和持续适应方面存在明显不足,而这些正是高效完成复杂任务的关键。本文揭示了当前LLM-MAS的优势与局限,并提供了统一、开源的评测平台以推动未来改进与评估。环境、30个开放任务及评估工具已公开于https://github.com/YusaeMeow/Collab-Overcooked。
原文摘要 · Abstract (English)
Large Language Models (LLMs) based agent systems have made great strides in real-world applications beyond traditional NLP tasks. This paper proposes a new LLM-based Multi-Agent System (LLM-MAS) benchmark, Collab-Overcooked, built on the popular Overcooked-AI game with more applicable and challenging tasks in interactive environments. Collab-Overcooked extends existing benchmarks in two novel ways. First, it provides a multi-agent framework supporting diverse tasks and objectives and encourages collaboration through natural language communication. Second, it introduces a spectrum of process-oriented evaluation metrics to assess the fine-grained collaboration capabilities of different LLM agents, a dimension often overlooked in prior work. We conduct extensive experiments with 13 popular LLMs and show that, while the LLMs exhibit a strong ability in goal interpretation, there are significant shortcomings in active collaboration and continuous adaptation, which are critical for efficiently fulfilling complex tasks. Notably, we highlight the strengths and weaknesses of LLM-MAS and provide insights for improving and evaluating LLM-MAS on a unified and open-source benchmark. The environments, 30 open-ended tasks, and the evaluation package are publicly available at https://github.com/YusaeMeow/Collab-Overcooked.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。