用智能评估代理提升深度研究模型的可信度,解决传统评测易被表面流畅性误导的问题。
DREAM: Deep Research Evaluation with Agentic Metrics
- 让评估过程本身具备智能代理能力,动态调用工具进行验证
- 在事实性和时间有效性上检测灵敏度远超现有基准
- 适合需要高可靠性评测的研究型AI系统开发者
深度研究智能体可生成分析师级报告,但其评估因缺乏统一标准和研究质量的多维特性而困难。现有基准存在‘合成幻觉’问题——表面流畅与引用对齐可能掩盖深层事实与推理错误。我们提出四维度分类框架揭示关键能力错配:静态评估者缺乏工具使用能力,无法判断时间有效性与事实正确性。为此,我们提出DREAM(基于智能评估指标的深度研究评估),通过让评估过程具备代理能力实现能力对等。DREAM采用查询无关指标与由调用工具的代理生成的自适应指标相结合的评估协议,支持时序感知覆盖、事实根植验证与系统性推理探测。受控实验表明,DREAM在识别事实性与时间性退化方面显著更敏感,提供可扩展、无需参考的评估范式。
原文摘要 · Abstract (English)
Deep Research Agents generate analyst-grade reports, yet evaluating them remains challenging due to the absence of a single ground truth and the multidimensional nature of research quality. Recent benchmarks propose distinct methodologies, yet they suffer from the Mirage of Synthesis, where strong surface-level fluency and citation alignment can obscure underlying factual and reasoning defects. We characterize this gap by introducing a taxonomy across four verticals that exposes a critical capability mismatch: static evaluators inherently lack the tool-use capabilities required to assess temporal validity and factual correctness. To address this, we propose DREAM (Deep Research Evaluation with Agentic Metrics), a framework that instantiates the principle of capability parity by making evaluation itself agentic. DREAM structures assessment through an evaluation protocol combining query-agnostic metrics with adaptive metrics generated by a tool-calling agent, enabling temporally aware coverage, grounded verification, and systematic reasoning probes. Controlled evaluations demonstrate DREAM is significantly more sensitive to factual and temporal decay than existing benchmarks, offering a scalable, reference-free evaluation paradigm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。