arXiv:2601.09688cs.CL2026-01被引 12

自动化构建复杂研究任务并智能评估大模型的深度研究能力

DeepResearchEval: An Automated Framework for Deep Research Task Construction and Agentic Evaluation

  • 基于用户角色生成真实复杂的多步研究任务
  • 动态生成评价维度,自动验证报告事实真伪
  • 适合评估大模型在无引用情况下的研究可靠性

深度研究系统广泛应用于多步网络调研、分析与跨源整合,但其评估仍具挑战。现有基准常需密集标注、依赖静态评价维度,或在缺乏引用时无法可靠验证事实。为此,我们提出 DeepResearchEval,一个自动化深度研究任务构建与代理评估框架。任务构建方面,采用角色驱动的流水线,生成基于多样用户画像的真实复杂任务,并通过两阶段筛选(任务资格与搜索必要性)保留需多源证据整合与外部检索的任务。评估方面,提出包含两个组件的代理流水线:自适应逐点质量评估,根据生成任务动态推导特定评价维度、标准与权重;主动事实核查,即使在无引用情况下也能自主提取并通过网络搜索验证报告陈述。

原文摘要 · Abstract (English)

Deep research systems are widely used for multi-step web research, analysis, and cross-source synthesis, yet their evaluation remains challenging. Existing benchmarks often require annotation-intensive task construction, rely on static evaluation dimensions, or fail to reliably verify facts when citations are missing. To bridge these gaps, we introduce DeepResearchEval, an automated framework for deep research task construction and agentic evaluation. For task construction, we propose a persona-driven pipeline generating realistic, complex research tasks anchored in diverse user profiles, applying a two-stage filter Task Qualification and Search Necessity to retain only tasks requiring multi-source evidence integration and external retrieval. For evaluation, we propose an agentic pipeline with two components: an Adaptive Point-wise Quality Evaluation that dynamically derives task-specific evaluation dimensions, criteria, and weights conditioned on each generated task, and an Active Fact-Checking that autonomously extracts and verifies report statements via web search, even when citations are missing.

深度研究自动评估代理系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。