arXiv:2506.06287cs.AI2025-06被引 31

评测AI网页研究代理性能,突破网页内容动态变化的难题。

Deep Research Bench: Evaluating AI Web Research Agents

  • 构建89个跨领域的多步研究任务,用人工精标答案确保评估可信。
  • 提出离线'回溯搜索'环境,使模型评测不受网页实时变化影响。
  • 可追踪幻觉、工具使用与遗忘行为,适合评估主流研究型AI产品。

当前最普遍的AI应用之一是具备网络搜索能力的大语言模型对话。然而,缺乏在不断变化的网络环境下对网络研究代理质量的直接评估。我们提出Deep Research Bench,包含89个跨8类主题、难度各异的多步网络研究任务实例,并由熟练人类精心标注答案。我们提供一个包含大量已冻结抓取网页的“RetroSearch”环境,证明离线‘回溯搜索’代理的表现可媲美在线实时搜索代理,实现对模型随时间演化的可靠评估。我们还提供完整的代理工具链和框架,用于持续评测新发布的主流LLM,包括o3和Gemini 2.5 Pro等具备推理能力的模型。通过自动化分析长时间的代理行为轨迹,可报告幻觉、工具调用及遗忘行为的变化趋势。最后,我们对以“Deep Research”“Deep Search”“Search”或“Research”命名的主要研究类产品进行了评估,结果公开于https://drb.futuresearch.ai/。

原文摘要 · Abstract (English)

Amongst the most common use cases of modern AI is LLM chat with web search enabled. However, no direct evaluations of the quality of web research agents exist that control for the continually-changing web. We introduce Deep Research Bench, consisting of 89 multi-step web research task instances of varying difficulty across 8 diverse task categories, with the answers carefully worked out by skilled humans. We provide a "RetroSearch" environment with a large frozen set of scraped web pages, and demonstrate that offline "RetroSearch" agents perform comparably to "live web" agents, enabling reliable evaluations of models over time. We provide robust agent tooling and scaffolding to benchmark major LLMs as they are released, including "thinking" models like o3 and Gemini 2.5 Pro. We include automated evaluations of the lengthy agent traces to report progress over time in hallucinations, tool use, and forgetting. Finally, we evaluate the major web research products branded as "Deep Research", "Deep Search", "Search", or "Research." Results are available on a public leaderboard at https://drb.futuresearch.ai/.

AI评测网页研究大模型评估幻觉检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。