arXiv:2508.15804cs.CLcs.AI2025-08综述被引 34

评测大模型生成学术报告的质量,发现商用代理仍存事实误差。

ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks

  • 用论文反向生成测试题,构建真实研究场景的评估集。
  • 商用代理报告比普通搜索模型更全面可靠,但仍有事实错误。
  • 适合研究大模型在科研中的可靠性,或做智能写作工具评估。

深度研究代理的出现显著缩短了开展广泛研究任务的时间。然而,这些任务本身对事实准确性和全面性要求极高,亟需严格的评估体系以支撑其广泛应用。本文提出ReportBench,一个系统性的基准测试框架,用于评估大语言模型(LLMs)生成的研究报告内容质量。评估聚焦两个核心维度:(1) 引用文献的质量与相关性,(2) 报告中陈述内容的真实性与忠实度。ReportBench利用arXiv上高质量的综述论文作为黄金标准参考,通过反向提示工程生成特定领域提示,并建立完整的评估语料库。此外,我们设计了一个基于代理的自动化分析框架,能系统提取报告中的引用与陈述,验证引用内容是否与原始文献一致,并使用网络资源验证未引用的主张。实证评估表明,如OpenAI和Google开发的商用深度研究代理,其生成报告在全面性和可靠性上优于仅靠搜索或浏览增强的独立大模型。然而,在研究覆盖广度与深度以及事实一致性方面仍存在较大改进空间。完整代码与数据将公开于:https://github.com/ByteDance-BandAI/ReportBench

原文摘要 · Abstract (English)

The advent of Deep Research agents has substantially reduced the time required for conducting extensive research tasks. However, these tasks inherently demand rigorous standards of factual accuracy and comprehensiveness, necessitating thorough evaluation before widespread adoption. In this paper, we propose ReportBench, a systematic benchmark designed to evaluate the content quality of research reports generated by large language models (LLMs). Our evaluation focuses on two critical dimensions: (1) the quality and relevance of cited literature, and (2) the faithfulness and veracity of the statements within the generated reports. ReportBench leverages high-quality published survey papers available on arXiv as gold-standard references, from which we apply reverse prompt engineering to derive domain-specific prompts and establish a comprehensive evaluation corpus. Furthermore, we develop an agent-based automated framework within ReportBench that systematically analyzes generated reports by extracting citations and statements, checking the faithfulness of cited content against original sources, and validating non-cited claims using web-based resources. Empirical evaluations demonstrate that commercial Deep Research agents such as those developed by OpenAI and Google consistently generate more comprehensive and reliable reports than standalone LLMs augmented with search or browsing tools. However, there remains substantial room for improvement in terms of the breadth and depth of research coverage, as well as factual consistency. The complete code and data will be released at the following link: https://github.com/ByteDance-BandAI/ReportBench

大模型评测研究代理报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。