arXiv:2510.14240cs.AI2025-10中稿 · ICLR被引 35

打造真实动态的深度研究评测基准,推动智能系统解决复杂信息整合难题。

LiveResearchBench: A Live Benchmark for User-Centric Deep Research in the Wild

  • 构建100个贴近真实需求的动态研究任务,需实时跨源搜索与综合分析。
  • 耗时超1500小时人工标注,确保任务可比性与评估可靠性。
  • 提出多维度评估框架DeepEval,精准衡量报告内容质量与引用准确性。

深度研究——通过从数百个实时网络来源检索并整合信息,生成全面且有引文支持的报告——是智能体系统的重要前沿。为严格评估该能力,必须满足四个原则:任务应以用户为中心、反映真实信息需求;具有动态性,要求超越参数化知识的最新信息;表述明确,避免歧义;且多维度、强搜索依赖,需跨大量网页进行深入分析。现有基准在这些方面存在不足,常局限于窄域或问题模糊,难以公平比较。基于上述原则,我们推出LiveResearchBench,包含100个专家精心设计的任务,覆盖日常生活、企业及学术领域,每项均需广泛、动态、实时的网络搜索与综合。该基准投入超过1500小时人力构建,为系统评估提供坚实基础。为评估有引文支撑的长篇报告,我们引入DeepEval,一套涵盖内容与报告层级的综合评估套件,包括覆盖率、呈现效果、引文准确性与关联性、分析一致性与深度。DeepEval整合四种互补评估协议,确保评估稳定且与人工判断高度一致。利用LiveResearchBench与DeepEval,我们对17个前沿深度研究系统(包括单智能体网页搜索、单智能体深度研究及多智能体系统)进行了全面评估。分析揭示了当前系统的优劣、常见失败模式,以及实现可靠、深刻深度研究的关键组件。代码已开源:https://github.com/SalesforceAIResearch/LiveResearchBench。

原文摘要 · Abstract (English)

Deep research -- producing comprehensive, citation-grounded reports by searching and synthesizing information from hundreds of live web sources -- marks an important frontier for agentic systems. To rigorously evaluate this ability, four principles are essential: tasks should be (1) user-centric, reflecting realistic information needs, (2) dynamic, requiring up-to-date information beyond parametric knowledge, (3) unambiguous, ensuring consistent interpretation across users, and (4) multi-faceted and search-intensive, requiring search over numerous web sources and in-depth analysis. Existing benchmarks fall short of these principles, often focusing on narrow domains or posing ambiguous questions that hinder fair comparison. Guided by these principles, we introduce LiveResearchBench, a benchmark of 100 expert-curated tasks spanning daily life, enterprise, and academia, each requiring extensive, dynamic, real-time web search and synthesis. Built with over 1,500 hours of human labor, LiveResearchBench provides a rigorous basis for systematic evaluation. To evaluate citation-grounded long-form reports, we introduce DeepEval, a comprehensive suite covering both content- and report-level quality, including coverage, presentation, citation accuracy and association, consistency and depth of analysis. DeepEval integrates four complementary evaluation protocols, each designed to ensure stable assessment and high agreement with human judgments. Using LiveResearchBench and DeepEval, we conduct a comprehensive evaluation of 17 frontier deep research systems, including single-agent web search, single-agent deep research, and multi-agent systems. Our analysis reveals current strengths, recurring failure modes, and key system components needed to advance reliable, insightful deep research. Our code is available at: https://github.com/SalesforceAIResearch/LiveResearchBench.

深度研究智能体评测动态基准引文评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。