arXiv:2508.07999cs.CL2025-08被引 57

测试智能体大规模信息搜集能力,发现现有系统表现极差。

WideSearch: Benchmarking Agentic Broad Info-Seeking

  • 构建200个真实用户问题的评测集,覆盖15个领域
  • 多数智能体成功率接近0%,最优仅达5%
  • 适合研究智能体搜索可靠性与改进方向的人看

从专业研究到日常规划,许多任务受限于大规模信息搜集,其难点在于重复性高而非认知复杂。随着大语言模型(LLMs)的发展,基于LLM的自动化搜索智能体有望解放人类完成此类繁琐工作。然而,当前缺乏合适的基准来评估这些智能体在“宽上下文”信息收集任务中的可靠性与完整性。为此,我们提出WideSearch,一个专为评估智能体在大规模信息收集任务中表现而设计的新基准。该基准包含200个手工标注的问题(100个英文,100个中文),覆盖超过15个不同领域,均源于真实用户查询。每个任务要求智能体收集大量可逐项验证的原子信息,并组织成结构化输出。通过严格的五阶段质量控制流程,确保数据集的难度、完整性和可验证性。我们对十余种先进代理搜索系统进行了评测,包括单智能体、多智能体框架及端到端商业系统。结果显示,大多数系统整体成功率接近0%,最优者仅达5%。但若给予充足时间并由多位人工测试者交叉验证,成功率可接近100%。这表明当前搜索智能体在大规模信息搜集方面存在严重缺陷,亟需未来研究突破。数据集、评估流程与评测结果已公开:https://widesearch-seed.github.io/

原文摘要 · Abstract (English)

From professional research to everyday planning, many tasks are bottlenecked by wide-scale information seeking, which is more repetitive than cognitively complex. With the rapid development of Large Language Models (LLMs), automated search agents powered by LLMs offer a promising solution to liberate humans from this tedious work. However, the capability of these agents to perform such "wide-context" collection reliably and completely remains largely unevaluated due to a lack of suitable benchmarks. To bridge this gap, we introduce WideSearch, a new benchmark engineered to evaluate agent reliability on these large-scale collection tasks. The benchmark features 200 manually curated questions (100 in English, 100 in Chinese) from over 15 diverse domains, grounded in real user queries. Each task requires agents to collect large-scale atomic information, which could be verified one by one objectively, and arrange it into a well-organized output. A rigorous five-stage quality control pipeline ensures the difficulty, completeness, and verifiability of the dataset. We benchmark over 10 state-of-the-art agentic search systems, including single-agent, multi-agent frameworks, and end-to-end commercial systems. Most systems achieve overall success rates near 0\%, with the best performer reaching just 5\%. However, given sufficient time, cross-validation by multiple human testers can achieve a near 100\% success rate. These results demonstrate that present search agents have critical deficiencies in large-scale information seeking, underscoring urgent areas for future research and development in agentic search. Our dataset, evaluation pipeline, and benchmark results have been publicly released at https://widesearch-seed.github.io/

智能体搜索信息搜集基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。