arXiv:2601.20975cs.CL2026-01被引 34

评测大模型深度搜索能力,发现现有系统召回与精度难兼顾。

DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents

  • 设计900个跨领域多步任务,考验复杂搜索规划能力
  • 顶尖模型在高召回下精度仍不足,普遍存在过早停止或滥搜现象
  • 适合研究深度搜索代理、信息检索与推理的学者使用

我们提出DeepSearchQA,一个包含900个提示的基准测试,用于评估智能体在17个不同领域中执行复杂多步信息查询任务的能力。与传统仅关注单答案检索或广谱事实性验证的基准不同,DeepSearchQA通过精心设计的挑战性任务,评估智能体系统性整合分散来源信息、去重与实体消歧、以及在开放搜索空间中判断终止条件的能力。每个任务以因果链形式组织,前一步的成功依赖于上一步完成,强调长程规划与上下文保持。所有任务均基于开放网络,答案集可客观验证。对当前最先进代理架构的全面评估显示显著性能瓶颈:即使最优模型也难以同时实现高召回与高精度。观察到的失败模式包括过早停止(召回不足)和过度泛化行为——为人为提升召回而生成大量低置信度答案。这些发现揭示了现有代理设计的巨大改进空间,并确立DeepSearchQA作为推动深度研究能力发展的关键诊断工具。

原文摘要 · Abstract (English)

We introduce DeepSearchQA, a 900-prompt benchmark for evaluating agents on difficult multi-step information-seeking tasks across 17 different fields. Unlike traditional benchmarks that target single answer retrieval or broad-spectrum factuality, DeepSearchQA features a dataset of challenging, handcrafted tasks designed to evaluate an agent's ability to execute complex search plans to generate exhaustive answer lists. This shift in design explicitly tests three critical, yet under-evaluated capabilities: 1) systematic collation of fragmented information from disparate sources, 2) de-duplication and entity resolution to ensure precision, and 3) the ability to reason about stopping criteria within an open-ended search space. Each task is structured as a causal chain, where discovering information for one step is dependent on the successful completion of the previous one, stressing long-horizon planning and context retention. All tasks are grounded in the open web with objectively verifiable answer sets. Our comprehensive evaluation of state-of-the-art agent architectures reveals significant performance limitations: even the most advanced models struggle to balance high recall with precision. We observe distinct failure modes ranging from premature stopping (under-retrieval) to hedging behaviors, where agents cast an overly wide net of low-confidence answers to artificially boost recall. These findings highlight critical headroom in current agent designs and position DeepSearchQA as an essential diagnostic tool for driving future research toward more robust, deep-research capabilities.

智能体深度搜索评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。