arXiv:2512.01948cs.CL2025-12被引 10

提出新基准与故障分类,揭示深度研究代理的三大核心短板。

How Far Are We from Genuinely Useful Deep Research Agents?

  • 构建100个真实研究任务与419项结构化检查清单,统一报告标准。
  • 分析1000份报告发现,现有代理在证据整合与验证上严重不足。
  • 首次系统梳理14类故障模式,适合评估与改进研究型AI代理。

深度研究代理(DRAs)旨在通过迭代的信息检索与综合生成分析师级报告。然而,现有大多数DRAs仅在问答基准上验证,对生成综合性报告的研究被忽视。更严重的是,当前报告合成基准存在任务复杂性高和主观评价指标问题,无法反映用户真实需求,限制了生成报告的实际效用。为解决这些缺口,我们提出了细粒度深度研究基准FINDER,包含100个人工精心策划的研究任务和419项结构化检查项,用于标准化报告结构、分析深度与事实依据。基于约1000份主流DRAs生成的报告,我们进一步提出深度研究失败分类法DEFT,这是首个针对深度研究代理的失败分类体系。DEFT包含14种细粒度失败模式,覆盖推理、检索与生成环节,基于扎根理论,经人类与大模型协同标注及标注者间一致性验证。实验发现,当前DRAs并非不理解任务,而是在证据整合、验证及推理稳健规划方面表现不佳。

原文摘要 · Abstract (English)

Deep Research Agents (DRAs) aim to automatically produce analyst-level reports through iterative information retrieval and synthesis. However, most existing DRAs were validated on question-answering benchmarks, while research on generating comprehensive reports remains overlooked. Worse, current benchmarks for report synthesis suffer from task complexity and subjective metrics -- this fails to reflect user demands and limits the practical utility of generated reports. To address these gaps, we present Fine-grained DEepResearch bench (FINDER), an enhanced benchmark consisting of 100 human-curated research tasks with 419 structured checklist items that standardize report structure, analytical depth, and factual grounding. Based on approximately 1,000 reports produced by mainstream DRAs, we further propose Deep rEsearch Failure Taxonomy (DEFT), the first failure taxonomy for deep research agents. DEFT contains 14 fine-grained failure modes across reasoning, retrieval, and generation, and is built upon grounded theory with human-LLM co-annotating and inter-annotator reliability validation. Our experimental findings reveal that current DRAs struggle not with task comprehension but with evidence integration, verification, and reasoning-resilient planning.

研究代理评估基准故障分析大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。