arXiv:2605.21482cs.AI2026-05

新基准挑战大模型深度推理能力,揭示真实短板。

DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation

论文配图:DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
图 1 · 摘自论文原文
  • 设计需跨源证据整合与长程推导的任务,检验深层推理能力。
  • 70%以上错误源于推导与校准不足,非检索问题。
  • 模型在不同领域表现差异大,存在真实能力专长。

深度研究指智能体在开放网络中搜索、收集证据,并通过多步推理得出答案。现有基准难以区分前沿模型的真实能力。我们提出DeepWeb-Bench,其难度显著高于当前主流基准,任务需大规模证据收集、跨源信息对齐及长程多步推导。我们将难点归纳为四类能力:检索、推导、推理与校准,并分项报告结果。每个答案附带四级来源溯源记录与跨源验证,提升可审计性。我们在九个前沿模型上评估,发现:(1) 检索失败仅占错误的12%-14%,而推导与校准错误占比超70%;(2) 强模型错误以推导不全为主,弱模型则多为虚构精确信息;(3) 模型在不同领域表现分化明显,跨模型一致率仅为rho = 0.61,单任务分歧达18.8个百分点。公开版本包含数据、评分标准与评测代码。

原文摘要 · Abstract (English)

Deep research, in which an agent searches the open web, collects evidence, and derives an answer through extended reasoning, is a prominent use case for frontier language models. Frontier deep research products score high on existing benchmarks, making it difficult to distinguish their capabilities from current evaluation data alone. We introduce DeepWeb-Bench, a deep research benchmark that is substantially harder than existing benchmarks for the current frontier. Difficulty comes from three properties of the data itself: each task requires massive evidence collection, cross-source reconciliation, and long-horizon multi-step derivation. We represent these three sources of difficulty as four capability families (Retrieval, Derivation, Reasoning, and Calibration) and report results sliced by family. Every reference answer is accompanied by a source-provenance record with four disclosure levels and cross-source checks where available, making scores easier to audit against the underlying evidence. We evaluate DeepWeb-Bench on nine frontier models and report three findings: (1) retrieval is not the bottleneck, as retrieval failures account for only 12-14% of errors while derivation and calibration failures account for over 70%; (2) strong and weak models fail in qualitatively different ways, with strong models' errors dominated by incomplete derivation and weak models' by hallucinated precision; and (3) models exhibit genuine specialization across domains, with cross-model agreement of only rho = 0.61 and per-case disagreement reaching 18.8 percentage points. The public benchmark release includes the data, rubrics, and evaluation code.

深度推理模型评估证据整合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。