对比多智能体与单智能体在统一协议下的表现,发现多数多智能体更差。
Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows

- 统一执行与日志协议,公平比较单智能体与多智能体流程。
- 六种多智能体中仅一种优于单智能体,其余低2.56至11.29分。
- 动态生成工作流在GAIA上超基线20分以上,适合复杂任务研究者。
在共享基准加载器、工具访问、答案合约、使用计费和轨迹日志的前提下,增加智能体是否有助于大语言模型工作流?我们提出BenchAgent评估框架,将单智能体、固定多智能体(MAS)及演化多智能体流程置于同一标准化执行与日志协议下。BenchAgent在十项推理、编码与工具使用基准上对GPT-4.1进行评估,并单独报告了运行时生成工作流的协议对齐外部(PAE)GAIA研究。在相同初始化(SI)条件下,六种测试的MAS中最多只有一种超过对应单智能体基准:EvoAgent处于威尔逊单次运行置信区间内,其余五种均落后2.56至11.29分,且代价更高。在PAE GAIA快照中,类似Claude-Code的运行时工作流整体达到66.72%,三级任务达69.23%,比最强非Claude基线Jarvis高出20多分。
原文摘要 · Abstract (English)
Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and trajectory logging? We introduce BenchAgent, an evaluation framework that places single-agent, fixed multi-agent (MAS), and evolving MAS workflows under one normalized execution and logging protocol. BenchAgent evaluates these substrate-internal workflows across ten reasoning, coding, and tool-use benchmarks with GPT-4.1, and separately reports a Protocol-Aligned External (PAE) GAIA study of a runtime-generated workflow. Under SI conditions, at most one of six tested MAS exceeds the matched single-agent anchor on benchmark-balanced average accuracy: EvoAgent lies within the Wilson one-run guidance, while the remaining five trail by 2.56-11.29 points and occupy more expensive accuracy-cost trade-offs. On the PAE GAIA snapshot, a Claude-Code-style runtime workflow reaches 66.72% overall and 69.23% on Level 3, more than 20 points above the strongest non-Claude baseline, Jarvis, a fixed MAS.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。