arXiv:2507.16280cs.AI2025-07被引 34

首个评估深度科研AI在前沿科学问题上发现新见解能力的基准

ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry

  • 构建涵盖65个真实科研场景的前沿问题数据集
  • 领先系统在开放咨询类问题上表现显著优于其他模型
  • 通过双评估框架量化洞察质量与事实准确性,推动科研协作新范式

深度研究系统在解决问题方面展现出强大能力,从基础查询扩展到复杂科研任务。然而,现有基准多聚焦于网络检索与报告生成,忽视其在科学前沿发现新见解的潜力。为此,我们提出ResearcherBench,首个针对深度人工智能科研系统(DARS)在前沿科学问题上能力评估的基准。数据集包含65个来自实验室讨论与访谈的真实科研问题,覆盖35个不同AI领域,分为技术细节、文献综述和开放咨询三类。采用双评估框架:基于专家设计标准的评阅评估(衡量洞察质量),以及基于引用准确性和覆盖面的事实评估(衡量忠实度与根基性)。我们评估了多个主流商业DARS及基线系统,结果显示OpenAI Deep Research和Gemini Deep Research在开放咨询类问题上显著领先。这些能力标志着向人工智能自我提升迈出关键一步,契合超级智能(ASI)愿景。我们开源ResearcherBench,为下一代科研助手的发展提供标准化平台,推动科学协作新范式:https://github.com/GAIR-NLP/ResearcherBench。

原文摘要 · Abstract (English)

The emergence of deep research systems presents significant capabilities in problem-solving, extending from basic queries to sophisticated research tasks. However, existing benchmarks primarily evaluate these systems as agents for web retrieval and report generation, overlooking their potential to discover novel insights on the frontiers of scientific research. To address this gap, we introduce ResearcherBench, the first benchmark focused on evaluating the capabilities of these advanced, agentic systems - which we refer to as Deep AI Research Systems (DARS) - on frontier AI scientific questions. We compiled a dataset of 65 research questions expertly selected from real-world scientific scenarios such as laboratory discussions and interviews, spanning 35 different AI subjects and categorized into three types: technical details, literature review, and open consulting. Our dual evaluation framework combines rubric assessment, which uses expert-designed criteria to evaluate insight quality, with factual assessment, which measures citation accuracy (faithfulness) and coverage (groundedness). We evaluated several leading commercial DARS and baseline systems. Results show that OpenAI Deep Research and Gemini Deep Research significantly outperform other systems, with particular strength in open-ended consulting questions. Such capabilities represent a meaningful step toward AI self-improvement, aligning with the vision of ASI for AI. We open-source ResearcherBench to provide a standardized platform for promoting the development of next-generation AI research assistants, hoping to foster a new perspective in AI research evaluation for a novel pattern of scientific collaboration: https://github.com/GAIR-NLP/ResearcherBench.

科研助手AI评估前沿探索

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。