首个评测搜索代理深度与广度结合能力的基准,揭示当前模型严重不足。
DeepWideSearch: Benchmarking Depth and Width in Agentic Information Seeking
- 构建新基准,要求代理同时完成多跳推理与大规模信息收集。
- 顶尖模型平均成功率仅2.39%,凸显整合难度。
- 暴露四大缺陷,适合研究智能搜索与代理系统的学者。
当前搜索代理在多跳检索的深度推理与大规模信息收集的广度之间存在根本性缺失,制约了其在市场分析等真实场景的应用。为此,我们提出DeepWideSearch,首个专门用于评估代理在信息寻求中融合深度与广度能力的基准。该基准包含220个问题,覆盖15个不同领域,每个问题需处理大量数据并进行多跳推理。实验表明,即使是最先进的代理,平均成功率为2.39%,凸显融合深度与广度的巨大挑战。进一步分析揭示四种失败模式:缺乏反思、过度依赖内部知识、检索不足和上下文溢出,暴露出当前代理架构的关键局限。我们公开发布DeepWideSearch,以推动更强大、更鲁棒的信息寻求代理研究。
原文摘要 · Abstract (English)
Current search agents fundamentally lack the ability to simultaneously perform \textit{deep} reasoning over multi-hop retrieval and \textit{wide}-scale information collection-a critical deficiency for real-world applications like comprehensive market analysis and business development. To bridge this gap, we introduce DeepWideSearch, the first benchmark explicitly designed to evaluate agents to integrate depth and width in information seeking. In DeepWideSearch, agents must process a large volume of data, each requiring deep reasoning over multi-hop retrieval paths. Specifically, we propose two methods to converse established datasets, resulting in a curated collection of 220 questions spanning 15 diverse domains. Extensive experiments demonstrate that even state-of-the-art agents achieve only 2.39% average success rate on DeepWideSearch, highlighting the substantial challenge of integrating depth and width search in information-seeking tasks. Furthermore, our error analysis reveals four failure modes: lack of reflection, overreliance on internal knowledge, insufficient retrieval, and context overflow-exposing key limitations in current agent architectures. We publicly release DeepWideSearch to catalyze future research on more capable and robust information-seeking agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。