评测大模型在深度研究中逐层整合证据的能力,发现报告质量高不等于推理可靠。
HiEviDR-Bench: A Benchmark for Hierarchical Evidence Aggregation in Deep Research

- 构建显式证据图,追踪从证据到结论的多阶段推理过程
- 2000个经人工验证的问题,覆盖文本与多模态场景,含不同难度层级
- 揭示多数模型在引用准确性和中间结论构建上表现差,表面报告质量不可靠
深度研究需要模型从大规模异构来源中检索、关联并综合证据以回答复杂问题并生成分析报告。现有基准主要评估最终结果,如答案正确性、报告质量或引用一致性,但难以观察证据是否被正确选择、链接和聚合为支持性论断与结论。为此,我们提出HiEviDR-Bench,一个用于评估深度研究中层次化证据聚合的基准。该基准涵盖开放域与学术域,在纯文本与多模态条件下均有覆盖,并以显式证据图表示每个实例,捕捉证据选取、跨源链接以及从证据到中间论断和最终结论的聚合过程。基于此,我们设计了以可追溯性为核心的五维评估框架:报告质量、证据可追溯性、引用准确性、论断验证和答案正确性,并引入渐进式门控机制实现细粒度错误定位。HiEviDR-Bench包含2000个经人工验证的问题及其证据图,覆盖多个难度等级。对16个代表性多模态大模型的实验表明,尽管许多系统在报告质量上表现良好,但在引用准确性、论断构建和答案正确性上显著下降。进一步分析显示,主要瓶颈在于证据识别和中间论断构建,说明强表面报告质量并不等同于在本基准上具备扎实的多阶段推理能力。
原文摘要 · Abstract (English)
Deep research requires models to retrieve, connect, and synthesize evidence from large-scale heterogeneous sources to answer complex queries and produce analytical reports. Existing benchmarks mainly evaluate final outcomes, such as answer correctness, report quality, or citation alignment, while providing limited visibility into whether evidence is correctly selected, linked, and aggregated into supported claims and conclusions. To address this gap, we introduce HiEviDR-Bench, a benchmark for evaluating Hierarchical Evidence Aggregation in Deep Research. HiEviDR-Bench covers open-domain and academic-domain settings under both text-only and multimodal conditions, and represents each instance with an explicit evidence graph that captures evidence selection, cross-source linking, and aggregation from evidence to intermediate claims and final conclusions. Based on this formulation, we develop a traceability-oriented evaluation framework with five dimensions: report quality, evidence traceability, citation accuracy, claim verification, and answer correctness, together with a progressive gating mechanism for fine-grained error localization. HiEviDR-Bench contains 2,000 human-validated questions with evidence graphs across multiple difficulty levels. Experiments on 16 representative multimodal large language models show that, although many systems achieve strong report quality, their performance drops markedly on citation accuracy, claim construction, and answer correctness. Further analysis shows that the main bottlenecks lie in evidence identification and intermediate claim construction, revealing that strong surface-level report quality does not necessarily imply grounded multi-stage reasoning on our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。