评测大模型生成专家级报告的综合基准,解决质量评估难题。
DEER: A Benchmark for Evaluating Deep Research Agents on Expert Report Generation
- 构建7维度25子维度的专家评级体系,含101个细粒度评分项。
- 提出声明验证架构,可检测引用与未引用陈述,量化证据质量。
- 揭示当前系统逻辑不完整问题,助力诊断改进方向。
大语言模型的发展推动了通过多步推理和基于证据合成生成专家级报告的深度研究系统。然而,报告质量评估面临多重挑战:质量维度复杂,难以确定评估标准;基于LLM的评判可能遗漏需领域知识才能识别的错误;且深度研究依赖检索证据,需进行报告层面的声明验证。为此,我们提出DEER基准,系统化评估标准,采用专家制定的7维25子维度分类体系,并转化为101个细粒度评分条目。同时提供任务特定的专家评估指南以支持基于LLM的评判。除基于评分表的评估外,还提出一种声明验证架构,能验证引用与未引用的声明,并量化证据质量。实验表明,当前系统虽生成结构合理、引用证据的报告,但仍难以完全满足专家级用户需求,逻辑完整性不足。DEER不仅可用于性能对比,更能使系统优劣可解释,提供可改进的诊断信号。
原文摘要 · Abstract (English)
Recent advances in large language models have enabled deep research systems that generate expert-level reports through multi-step reasoning and evidence-based synthesis. However, evaluating such reports remains challenging: report quality is multifaceted, making it difficult to determine what to assess and which criteria to use; LLM-based judges may miss errors that require domain expertise to identify; and because deep research relies on retrieved evidence, report-wide claim verification is also necessary. To address these issues, we propose DEER, a benchmark for evaluating expert-level deep research reports. DEER systematizes evaluation criteria with an expert-developed taxonomy (7 dimensions, 25 subdimensions) operationalized as 101 fine-grained rubric items. We also provide task-specific Expert Evaluation Guidance to support LLM-based judging. In addition to rubric-based assessment, we propose a claim verification architecture that verifies both cited and uncited claims and quantifies evidence quality. Experiments show that current systems produce structurally plausible, evidence-citing reports, but still struggle to fully satisfy expert-level user requests and achieve logical completeness. Beyond performance comparisons, DEER makes system strengths and limitations interpretable and provides diagnostic signals for improvement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。