评测大模型代理在协作与竞争中的表现,揭示最佳组织结构与策略。
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents

- 设计多代理互动场景,用里程碑指标评估协作与竞争质量。
- 图结构协作优于星型、链式等,认知规划提升任务达成率3%。
- 适合研究多智能体系统、具身智能或博弈机制的学者参考。
大型语言模型(LLMs)在作为自主代理方面展现出显著能力,但现有基准测试或聚焦单代理任务,或局限于特定领域,难以捕捉多代理协调与竞争的动态。本文提出MultiAgentBench,一个全面的基准框架,用于评估基于LLM的多代理系统在多样化交互场景中的表现。该框架不仅衡量任务完成度,还通过新颖的基于里程碑的关键性能指标,评估协作与竞争质量。我们评估了多种协调协议(包括星型、链式、树形和图结构)以及创新策略,如群体讨论与认知规划。结果显示,gpt-4o-mini在平均任务得分上最高,图结构在科研场景中表现最佳,认知规划使里程碑达成率提升3%。代码与数据集已公开于https://github.com/MultiagentBench/MARBLE。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents, yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario, and cognitive planning improves milestone achievement rates by 3%. Code and datasets are public available at https://github.com/MultiagentBench/MARBLE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。