评测AI在复杂科研文献发现中的自主能力,挑战远超现有基准。
AutoResearchBench: Benchmarking AI Agents on Complex Scientific Literature Discovery

- 设计双任务评测框架:深挖式追踪与广搜式收集。
- 顶尖大模型在深度研究上仅9.39%准确率,广搜任务IoU为9.31%。
- 专为科研场景设计,适合评估真实科研中AI的推理与搜索能力。
自主科学研宄因AI代理的发展而显著推进,其中关键步骤是发现合适的科学文献,无论是探索已有知识,还是为假设验证与主张提供证据。为评估AI代理在此过程中的能力,我们提出AutoResearchBench,一个专注于自主科研文献发现的专用基准。该基准包含两类互补任务:(1) 深度研究,要求通过多步渐进式探查定位特定目标论文;(2) 广度研究,要求全面收集满足给定条件的一组论文。相较于以往的代理网页浏览基准,AutoResearchBench在三个维度上具有独特性:研究导向,要求对科学概念有深入理解;文献聚焦,需精细利用详细信息;开放结尾,合格论文数量未知,需持续推理与搜索。这些特性使其特别适合评估自主研究能力,也极具挑战性。即使最强大的大语言模型,在深度研究任务上准确率也仅为9.39%,广度研究任务的IoU为9.31%,许多其他基线模型低于5%。我们已公开数据集、评估流程和代码,详见https://github.com/CherYou/AutoResearchBench。
原文摘要 · Abstract (English)
Autonomous scientific research is significantly advanced thanks to the development of AI agents. One key step in this process is finding the right scientific literature, whether to explore existing knowledge for a research problem, or to acquire evidence for verifying assumptions and supporting claims. To assess AI agents' capability in driving this process, we present AutoResearchBench, a dedicated benchmark for autonomous scientific literature discovery. AutoResearchBench consists of two complementary task types: (1) Deep Research, which requires tracking down a specific target paper through a progressive, multi-step probing process, and (2) Wide Research, which requires comprehensively collecting a set of papers satisfying given conditions. Compared to previous benchmarks on agentic web browsing, AutoResearchBench is distinguished along three dimensions: it is research-oriented, calling for in-depth comprehension of scientific concepts; literature-focused, demanding fine-grained utilization of detailed information; and open-ended, involving an unknown number of qualified papers and thus requiring deliberate reasoning and search throughout. These properties make AutoResearchBench uniquely suited for evaluating autonomous research capabilities, and extraordinarily challenging. Even the most powerful LLMs, despite having largely conquered general agentic web-browsing benchmarks such as BrowseComp, achieve only 9.39% accuracy on Deep Research and 9.31% IoU on Wide Research, while many other strong baselines fall below 5%. We publicly release the dataset and evaluation pipeline to facilitate future research in this direction. We publicly release the dataset, evaluation pipeline, and code at https://github.com/CherYou/AutoResearchBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。