arXiv:2508.10152cs.AI2025-08被引 2

开源研究代理性能提升至10%,推动开放AI研究发展

Improving and Evaluating Open Deep Research Agents

  • 提出改进版ODR+,通过三项策略优化搜索与推理能力
  • 在60题小规模基准上实现10%准确率,领先现有系统
  • 提供可复现的轻量级评测集BC-Small,适合学术研究

本文聚焦深度研究代理(DRAs),即能自主检索并利用网络内容回答自然语言问题的系统。尽管近期成果亮眼,但多数为闭源系统。目前仅有一个开源版本名为Open Deep Research(ODR)。我们针对挑战性榜单BrowseComp,提出更易计算的子集BC-Small,用于学术实验室评估。在该基准上,原始ODR与两个闭源系统(Anthropic、Google)均取得0%准确率。我们对ODR进行三项战略改进,形成ODR+模型,在60题测试集上达到10%的成功率,成为开闭源系统中的最佳表现。消融实验表明三项改进均有效。

原文摘要 · Abstract (English)

We focus here on Deep Research Agents (DRAs), which are systems that can take a natural language prompt from a user, and then autonomously search for, and utilize, internet-based content to address the prompt. Recent DRAs have demonstrated impressive capabilities on public benchmarks however, recent research largely involves proprietary closed-source systems. At the time of this work, we only found one open-source DRA, termed Open Deep Research (ODR). In this work we adapt the challenging recent BrowseComp benchmark to compare ODR to existing proprietary systems. We propose BrowseComp-Small (BC-Small), comprising a subset of BrowseComp, as a more computationally-tractable DRA benchmark for academic labs. We benchmark ODR and two other proprietary systems on BC-Small: one system from Anthropic and one system from Google. We find that all three systems achieve 0% accuracy on the test set of 60 questions. We introduce three strategic improvements to ODR, resulting in the ODR+ model, which achieves a state-of-the-art 10% success rate on BC-Small among both closed-source and open-source systems. We report ablation studies indicating that all three of our improvements contributed to the success of ODR+.

研究代理开源模型评测基准AI推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。