评测深度研究型AI的报告质量,揭示其真实研究能力。
Understanding DeepResearch via Reports
- 用研究报告评估AI研究能力,引入三维度量化框架。
- 在100个真实问题上测试4个商用系统,发现性能差异明显。
- 提供开源数据集和代码,支持系统对比与改进。
DeepResearch智能体代表一种变革性的人工智能范式,通过复杂推理与多工具协同实现专家级研究。然而,由于研究任务开放性强,且现有基准仅关注孤立能力,对其评估仍具挑战性。与传统大模型任务不同,DeepResearch需整合多元信息、生成洞见并输出连贯报告,这些能力难以简单验证。为此,我们提出DeepResearch-ReportEval,一个基于研究报告的综合性评估框架。该框架通过创新的LLM作为评判者方法,系统衡量报告的质量、冗余度与事实性,达到与专家高度一致的评估效果。我们构建了涵盖12类真实场景的100个精选查询标准基准,支持系统间的能力比较。对四个领先商业系统的评估揭示了各异的设计理念与性能权衡,为DeepResearch从信息助手向智能研究伙伴演进提供了基础洞察。源代码与数据已开源:https://github.com/HKUDS/DeepResearch-Eval。
原文摘要 · Abstract (English)
DeepResearch agents represent a transformative AI paradigm, conducting expert-level research through sophisticated reasoning and multi-tool integration. However, evaluating these systems remains critically challenging due to open-ended research scenarios and existing benchmarks that focus on isolated capabilities rather than holistic performance. Unlike traditional LLM tasks, DeepResearch systems must synthesize diverse sources, generate insights, and present coherent findings, which are capabilities that resist simple verification. To address this gap, we introduce DeepResearch-ReportEval, a comprehensive framework designed to assess DeepResearch systems through their most representative outputs: research reports. Our approach systematically measures three dimensions: quality, redundancy, and factuality, using an innovative LLM-as-a-Judge methodology achieving strong expert concordance. We contribute a standardized benchmark of 100 curated queries spanning 12 real-world categories, enabling systematic capability comparison. Our evaluation of four leading commercial systems reveals distinct design philosophies and performance trade-offs, establishing foundational insights as DeepResearch evolves from information assistants toward intelligent research partners. Source code and data are available at: https://github.com/HKUDS/DeepResearch-Eval.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。