arXiv:2506.15741cs.AIcs.CL2025-06EMNLP被引 27

系统评估代理设计对性能的影响,提出更可靠的评测方法。

OAgents: An Empirical Study of Building Effective Agents

  • 在GAIA和BrowseComp上公平测试关键组件设计影响
  • 发现部分看似合理的组件实为冗余,存在显著随机波动
  • 开源OAgents框架,模块化设计助力未来研究

近年来,代理型AI日益成为研究热点。然而我们指出,当前代理研究缺乏标准化与科学严谨性,导致方法间难以公平比较。因此,不同设计选择如何影响代理有效性尚不明确,进展衡量也面临挑战。本文在GAIA基准与BrowseComp上开展系统性实证研究,以公平严谨的方式考察主流代理框架中关键组件的设计影响。研究发现,缺乏标准评估协议导致以往工作(包括开源项目)不可复现,且随机运行间存在显著差异。为此,我们引入更稳健的评估协议以稳定比较。研究揭示哪些组件真正关键,而其他组件虽看似合理却属冗余。基于此,我们构建并开源了OAgents——一个在开源项目中达到顶尖性能的新基础代理框架,具备模块化设计,推动后续代理型AI研究。

原文摘要 · Abstract (English)

Recently, Agentic AI has become an increasingly popular research field. However, we argue that current agent research practices lack standardization and scientific rigor, making it hard to conduct fair comparisons among methods. As a result, it is still unclear how different design choices in agent frameworks affect effectiveness, and measuring their progress remains challenging. In this work, we conduct a systematic empirical study on GAIA benchmark and BrowseComp to examine the impact of popular design choices in key agent components in a fair and rigorous manner. We find that the lack of a standard evaluation protocol makes previous works, even open-sourced ones, non-reproducible, with significant variance between random runs. Therefore, we introduce a more robust evaluation protocol to stabilize comparisons. Our study reveals which components and designs are crucial for effective agents, while others are redundant, despite seeming logical. Based on our findings, we build and open-source OAgents, a new foundation agent framework that achieves state-of-the-art performance among open-source projects. OAgents offers a modular design for various agent components, promoting future research in Agentic AI.

代理系统实证研究可复现性开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。