提出新基准与方法,评估并提升大模型的多智能体协同能力
SwarmBench: Can Large Language Models Act as Agent Swarm Orchestrators?

- 构建多维度评测框架,从准确率、效率、成本和流程质量评估
- 发现现有模型在协同调度上差异显著,尤其在过程质量上表现不一
- 提出基于经验提取与回放的方法,稳定提升协同性能
基于大语言模型的多智能体系统正从固定交互结构演变为动态协同的智能体集群。然而,现有基准仍以单智能体或通用任务为主,难以系统评估关键的协调能力。本文提出SwarmBench,一个从准确率、效率、成本和过程质量多角度评估模型性能的基准。实验表明,当前模型在协调能力上存在显著差异,不仅体现在最终结果,更反映在协调过程的整体质量。基于此发现,我们进一步提出SwarmExp,一种基于经验提取与经验回放的简单而有效的方法,能持续提升大语言模型的协调表现。
原文摘要 · Abstract (English)
Large language model-based multi-agent systems are evolving from fixed interaction topologies toward dynamically orchestrated Agent Swarms. However, existing benchmarks are still largely based on single-agent or general-purpose agent tasks, making it difficult to systematically evaluate key orchestration capabilities. We propose SwarmBench, a benchmark that evaluates model performance from multiple perspectives, including accuracy, efficiency, cost, and process quality. Experimental results show that current models exhibit substantial differences in orchestration capability. These differences are reflected not only in final accuracy, efficiency, and cost, but also in the overall quality of the orchestration process itself. Based on these findings, we further propose SwarmExp, a simple yet effective method based on experience extraction and experience replay, which consistently improves the orchestration performance of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。