自动多智能体系统实际表现不如单智能体,昂贵却无效。
The Illusion of Multi-Agent Advantage

- 用自动生成的多智能体对比单智能体推理链,发现前者更慢更差。
- 在交互式多步任务中,自动多智能体性能全面落后于CoT-SC,成本高10倍。
- 揭示自动化设计导致架构臃肿,表面复杂无实际收益。
主流观点认为多智能体系统(MAS)优于单智能体系统(SAS),具备上下文保护、并行处理和分布式决策等优势。然而,现有实证支持主要基于仅强调孤立推理任务的基准测试,无法有效评估这些优势。本文针对自动生成的、旨在提升泛化能力的MAS,与链式思维自洽性(CoT-SC)进行系统性对比。在传统推理数据集及包含交互式多步流程的任务(如BrowseComp-Plus)上,结果显示自动生成的MAS始终低于CoT-SC,且成本高达10倍。为排除任务结构限制,我们设计了一个诊断性合成数据集,专门体现任务分解、上下文分离与并行潜力。结果表明,人工精心设计的MAS在该数据集上显著优于自动生成架构,兼具更高性能与成本效率。进一步分析显示,当前自动化设计产生大量冗余架构,表面复杂但功能无效,暴露了其与多智能体核心原则的根本错位。
原文摘要 · Abstract (English)
Prevailing wisdom posits that Multi-Agent Systems (MAS) are superior to Single-Agent Systems (SAS), citing advantages like context protection, parallel processing and distributed decision-making. However, empirical support for this claim relies primarily on comparisons with SAS baselines using benchmarks that prioritize isolated reasoning tasks, which do not adequately assess these advantages. Focusing on automatically generated MAS that are designed for enhanced generalizability over manually-designed counterparts, we perform a rigorous, systematic evaluation against SAS, specifically Chain-of-Thought with Self-Consistency (CoT-SC). Across traditional reasoning datasets and tasks with interactive multi-step workflows (e.g., BrowseComp-Plus), we demonstrate that automatic MAS consistently underperform CoT-SC despite being up to 10x more expensive. To isolate these failures from limitations inherent to task structure, we introduce a diagnostic synthetic dataset tailored for MAS featuring explicit task decomposition, context separation and parallelization potential. We show that expert-architected MAS consistently outperforms automatically generated architectures in both raw performance and cost-efficiency on this dataset, demonstrating that existing evaluation frameworks mask critical architectural gaps and inefficiencies of complex MAS by failing to account for the marginal utility of increased computational cost. Critically, systematic deconstruction of the generated MAS architectures reveals that current automated design paradigms produce architectural bloat that prioritizes superficial complexity which does not translate into functional utility, exposing a fundamental misalignment with multi-agent principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。