arXiv:2604.14683cs.AI2026-04

构建可复现的深度研究智能体评估基准,真实模拟科研全流程。

DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation

论文配图:DR$^{3}$-Eval: Towards Realistic and Reproducible Deep Research Evaluation
图 1 · 摘自论文原文
  • 基于真实用户材料与静态沙盒环境,还原复杂网络场景
  • 多维度评测信息召回、事实准确率等五项指标,与人工判断一致
  • 揭示当前模型在检索鲁棒性和幻觉控制上的严重缺陷

深度研究智能体(DRAs)旨在解决涉及规划、检索、多模态理解与报告生成的复杂长周期研究任务,但其评估因动态网络环境和模糊的任务定义而困难。我们提出 DR$^{3}$-Eval,一个面向多模态、多文件报告生成的现实且可复现的评估基准。该基准源自真实用户提供的材料,并配以每任务独立的静态研究沙盒语料库,模拟开放网络复杂性的同时保持完全可验证性,包含支持文档、干扰项和噪声。此外,我们引入多维度评估框架,测量信息召回率、事实准确性、引用覆盖率、指令遵循度和深度质量,并验证其与人工判断的一致性。基于多个先进语言模型开发的多智能体系统 DR$^{3}$-Agent 实验表明,DR$^{3}$-Eval 具有极高挑战性,暴露出检索鲁棒性和幻觉控制中的关键失败模式。代码与数据已公开。

原文摘要 · Abstract (English)

Deep Research Agents (DRAs) aim to solve complex, long-horizon research tasks involving planning, retrieval, multimodal understanding, and report generation, yet their evaluation remains challenging due to dynamic web environments and ambiguous task definitions. We propose DR$^{3}$-Eval, a realistic and reproducible benchmark for evaluating deep research agents on multimodal, multi-file report generation. DR$^{3}$-Eval is constructed from authentic user-provided materials and paired with a per-task static research sandbox corpus that simulates open-web complexity while remaining fully verifiable, containing supportive documents, distractors, and noise. Moreover, we introduce a multi-dimensional evaluation framework measuring Information Recall, Factual Accuracy, Citation Coverage, Instruction Following, and Depth Quality, and validate its alignment with human judgments. Experiments with our developed multi-agent system DR$^{3}$-Agent based on multiple state-of-the-art language models demonstrate that DR$^{3}$-Eval is highly challenging and reveals critical failure modes in retrieval robustness and hallucination control. Our code and data are publicly available.

智能体评估多模态可复现性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。