构建企业级智能体架构评测基准,揭示不同设计组合的性能差异。
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
- 设计18种智能体配置,测试四种核心架构维度的组合效果。
- 复杂任务最高成功率仅35.3%,简单任务达70.8%,整体表现有限。
- 发现模型偏好显著,提示未来设计需避免‘一刀切’方案。
尽管智能体架构的各个组件已被单独研究,但对复杂多智能体系统中不同设计维度如何相互作用仍缺乏充分实证理解。本研究通过构建一个面向企业场景的综合性基准,评估了基于主流大语言模型的18种不同智能体配置。我们考察了四个关键维度:编排策略、智能体提示实现方式(ReAct vs 函数调用)、记忆架构以及思维工具集成。实验结果揭示了显著的模型特定架构偏好,挑战了当前智能体系统中普遍存在的‘一刀切’范式。此外,企业在复杂任务上的整体智能体性能仍有明显短板,表现最佳的模型在复杂任务上最高仅达35.3%的成功率,在简单任务上为70.8%。这些发现旨在为未来智能体系统的设计提供实证依据,支持更科学的组件选择与模型决策。
原文摘要 · Abstract (English)
While individual components of agentic architectures have been studied in isolation, there remains limited empirical understanding of how different design dimensions interact within complex multi-agent systems. This study aims to address these gaps by providing a comprehensive enterprise-specific benchmark evaluating 18 distinct agentic configurations across state-of-the-art large language models. We examine four critical agentic system dimensions: orchestration strategy, agent prompt implementation (ReAct versus function calling), memory architecture, and thinking tool integration. Our benchmark reveals significant model-specific architectural preferences that challenge the prevalent one-size-fits-all paradigm in agentic AI systems. It also reveals significant weaknesses in overall agentic performance on enterprise tasks with the highest scoring models achieving a maximum of only 35.3\% success on the more complex task and 70.8\% on the simpler task. We hope these findings inform the design of future agentic systems by enabling more empirically backed decisions regarding architectural components and model selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。