arXiv:2508.04183cs.CL2025-08被引 28

定义深度研究任务并构建评估基准,揭示当前模型搜索能力的短板。

Characterizing Deep Research: A Benchmark and Formal Definition

  • 以概念广度和推理探索为核心,重新定义深度研究任务。
  • 100个科学与热点事件任务中,最佳模型F1仅0.55,多数低于0.3。
  • 适合关注复杂信息检索与推理能力提升的研究者使用。

撰写综述或分析报告等信息任务需要复杂的搜索与推理,近年来被统称为‘深度研究’(Deep Research, DR),也被新模型所采纳。然而,该任务的范围仍不明确,与其它推理密集型问题的区分尚不清晰。本文提出对深度研究任务的正式定义,并引入一个评估基准。我们认为,深度研究的核心特征并非生成长篇报告,而是搜索过程中所需的概念高并发探索,即广泛且深度的推理性探索。为实现客观评估,我们采用中间输出表示,编码搜索中发现的关键主张,从而将推理挑战与表面报告生成分离。基于此,我们构建了涵盖科学主题(如数据集、材料发现、先有技术检索)和公共兴趣事件(如航班事故、电影奖项)的多样化、高挑战性基准LiveDRBench,共包含100个任务。在最先进模型中,各子类别的F1得分介于0.02至0.72之间,OpenAI模型表现最优,整体F1为0.55。对推理轨迹的分析揭示了当前系统引用来源数量、分支与回溯行为的分布,为改进搜索机制与事实溯源能力指明方向。基准代码已开源:https://github.com/microsoft/LiveDRBench。

原文摘要 · Abstract (English)

Information tasks such as writing surveys or analytical reports require complex search and reasoning, and have recently been grouped under the umbrella of \textit{deep research} -- a term also adopted by recent models targeting these capabilities. Despite growing interest, the scope of the deep research task remains underdefined and its distinction from other reasoning-intensive problems is poorly understood. In this paper, we propose a formal characterization of the deep research (DR) task and introduce a benchmark to evaluate the performance of DR systems. We argue that the core defining feature of deep research is not the production of lengthy report-style outputs, but rather the high fan-out over concepts required during the search process, i.e., broad and reasoning-intensive exploration. To enable objective evaluation, we define DR using an intermediate output representation that encodes key claims uncovered during search-separating the reasoning challenge from surface-level report generation. Based on this formulation, we propose a diverse, challenging benchmark LiveDRBench with 100 challenging tasks over scientific topics (e.g., datasets, materials discovery, prior art search) and public interest events (e.g., flight incidents, movie awards). Across state-of-the-art DR systems, F1 score ranges between 0.02 and 0.72 for any sub-category. OpenAI's model performs the best with an overall F1 score of 0.55. Analysis of reasoning traces reveals the distribution over the number of referenced sources, branching, and backtracking events executed by current DR systems, motivating future directions for improving their search mechanisms and grounding capabilities. The benchmark is available at https://github.com/microsoft/LiveDRBench.

深度研究评估基准信息检索推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。