评测大模型研究代理的100个博士级任务基准
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents

- 用22个领域专家设计的100个任务评估研究代理
- 提出两种方法,精准匹配人类评判标准
- 开源基准与评估框架,助力智能研究工具发展
深度研究代理(Deep Research Agents, DRAs)是一类基于大语言模型的智能体,能自主完成多步骤网络探索、定向检索与高阶信息整合,将海量在线信息转化为分析师级别的、带引用的报告,将数小时的手动调研压缩至几分钟。然而,系统评估这类代理能力的全面基准仍缺失。为此,我们提出 DeepResearch Bench,包含100个由各领域专家精心设计的博士级研究任务,覆盖22个不同学科。评估DRAs极具挑战且耗时,因此我们提出了两种新方法:第一种为基于参考的评估方法,采用自适应标准判断生成报告质量;第二种框架用于评估代理的信息检索与收集能力,通过有效引用数量和引用准确率来衡量。我们已将 DeepResearch Bench 及核心评估框架开源至 https://github.com/Ayanami0730/deep_research_bench,以推动实用化大模型研究代理的发展。
原文摘要 · Abstract (English)
Deep Research Agents are a prominent category of LLM-based agents. By autonomously orchestrating multistep web exploration, targeted retrieval, and higher-order synthesis, they transform vast amounts of online information into analyst-grade, citation-rich reports--compressing hours of manual desk research into minutes. However, a comprehensive benchmark for systematically evaluating the capabilities of these agents remains absent. To bridge this gap, we present DeepResearch Bench, a benchmark consisting of 100 PhD-level research tasks, each meticulously crafted by domain experts across 22 distinct fields. Evaluating DRAs is inherently complex and labor-intensive. We therefore propose two novel methodologies that achieve strong alignment with human judgment. The first is a reference-based method with adaptive criteria to assess the quality of generated research reports. The other framework is introduced to evaluate DRA's information retrieval and collection capabilities by assessing its effective citation count and overall citation accuracy. We have open-sourced DeepResearch Bench and key components of these frameworks at https://github.com/Ayanami0730/deep_research_bench to accelerate the development of practical LLM-based agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。