arXiv:2510.00172cs.CL2025-10被引 14

构建企业级深度研究真实评测基准,检验AI Agent跨源信息整合能力。

DRBench: A Realistic Benchmark for Enterprise Deep Research

  • 设计多步骤任务合成流程,融合公开网络与私有知识库
  • 覆盖10个领域100个任务,评估报告准确性与逻辑连贯性
  • 适合评估企业AI代理在复杂场景下的实际表现

我们提出DRBench,一个面向企业场景下复杂开放式深度研究任务的评测基准。不同于以往聚焦简单问题或仅依赖网页查询的基准,DRBench要求智能体在多步查询(如“为确保符合该标准,应如何调整产品路线图?”)中,从公共网络和公司内部知识库中识别支持性事实。每项任务基于真实用户角色和企业背景,覆盖生产力软件、云文件系统、邮件、聊天记录及开放网络等异构数据源。任务通过人工参与验证的合成管道生成,评估指标包括相关洞察召回率、事实准确性及报告结构化程度。我们发布了涵盖销售、网络安全、合规等10个领域的100个深度研究任务。通过评估多种开源与闭源模型(如GPT、Llama、Qwen)及策略,揭示了当前AI代理在企业级深度研究中的优劣,指明了未来发展方向。代码与数据已开源:https://github.com/ServiceNow/drbench。

原文摘要 · Abstract (English)

We introduce DRBench, a benchmark for evaluating AI agents on complex, open-ended deep research tasks in enterprise settings. Unlike prior benchmarks that focus on simple questions or web-only queries, DRBench evaluates agents on multi-step queries (for example, "What changes should we make to our product roadmap to ensure compliance with this standard?") that require identifying supporting facts from both the public web and private company knowledge base. Each task is grounded in realistic user personas and enterprise context, spanning a heterogeneous search space that includes productivity software, cloud file systems, emails, chat conversations, and the open web. Tasks are generated through a carefully designed synthesis pipeline with human-in-the-loop verification, and agents are evaluated on their ability to recall relevant insights, maintain factual accuracy, and produce coherent, well-structured reports. We release 100 deep research tasks across 10 domains, such as Sales, Cybersecurity, and Compliance. We demonstrate the effectiveness of DRBench by evaluating diverse DR agents across open- and closed-source models (such as GPT, Llama, and Qwen) and DR strategies, highlighting their strengths, weaknesses, and the critical path for advancing enterprise deep research. Code and data are available at https://github.com/ServiceNow/drbench.

企业AI深度研究评测基准多源融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。