为深度研究智能体设计多维评估框架,全面检验其报告生成能力
Dr. Bench: A Multidimensional Evaluation for Deep Research Agents, from Answers to Reports
- 构建10大领域214个挑战任务,支持长篇报告式输出评估
- 引入语义质量、主题聚焦与检索可信度等多维度评分指标
- 适合研究智能体架构与复杂任务求解的开发者使用
作为向互联架构演进的智能体体现,深度研究智能体(DRAs)在任务分解、跨源检索、多阶段推理、信息整合和结构化输出方面表现出显著能力,显著提升复杂开放任务的表现。然而现有基准在评估维度、响应格式和评分机制上仍存在不足。本文提出Dr. Bench,一个专为DRAs及长篇报告式响应设计的多维评估框架。该基准包含10个广泛领域的214个专家精心设计的挑战任务,每项任务均配有手工构建的参考数据包,以支持综合评估。框架引入语义质量、主题聚焦度和检索可信度等指标,实现对DRAs生成长报告的全面评估。大量实验表明主流DRAs优于基于网络搜索增强的推理模型,但仍存在巨大改进空间。本研究为智能体能力评估、架构优化与范式演进提供了坚实基础。
原文摘要 · Abstract (English)
As an embodiment of intelligence evolution toward interconnected architectures, Deep Research Agents (DRAs) systematically exhibit the capabilities in task decomposition, cross-source retrieval, multi-stage reasoning, information integration, and structured output, which markedly enhance performance on complex and open-ended tasks. However, existing benchmarks remain deficient in evaluation dimensions, response format, and scoring mechanisms, limiting their effectiveness in assessing such agents. This paper introduces Dr. Bench, a multidimensional evaluation framework tailored to DRAs and long-form report-style responses. The benchmark comprises 214 expert-curated challenging tasks across 10 broad domains, each accompanied by manually constructed reference bundles to support composite evaluation. This framework incorporates metrics for semantic quality, topical focus, and retrieval trustworthiness, enabling a comprehensive evaluation of long reports generated by DRAs. Extensive experimentation confirms the superior performance of mainstream DRAs over web-search-tool-augmented reasoning models, yet reveals considerable scope for further improvement. This study provides a robust foundation for capability assessment, architectural refinement, and paradigm advancement of DRAs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。