arXiv:2508.15658cs.CLcs.AI2025-08综述被引 15

构建首个科学综述生成基准,评估大模型生成综述的能力

SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation

  • 设计包含主题、专家综述和引用的测试集
  • 覆盖百万级论文,量化生成综述的全面性与准确性
  • 适合研究自动化文献综述或评测大模型能力的学者

学术文献的快速膨胀使得人工撰写科学综述日益困难。尽管大语言模型在自动化这一过程方面展现出潜力,但该领域进展受限于缺乏标准化基准和评估协议。为此,我们提出SurGE(Survey Generation Evaluation),一个面向计算机科学领域的科学综述生成基准。SurGE包含两部分:(1) 测试实例集合,每条包含主题描述、专家撰写的综述及完整参考文献;(2) 超过一百万篇论文的大规模学术语料库。此外,我们还提出了一个自动评估框架,从四个维度衡量生成综述的质量:全面性、引用准确性、结构组织性和内容质量。对多种基于LLM的方法的评估揭示了显著性能差距,表明即使先进的代理型框架也难以应对综述生成的复杂性,凸显该领域未来研究的重要性。所有代码、数据与模型均已开源。

原文摘要 · Abstract (English)

The rapid growth of academic literature makes the manual creation of scientific surveys increasingly infeasible. While large language models show promise for automating this process, progress in this area is hindered by the absence of standardized benchmarks and evaluation protocols. To bridge this critical gap, we introduce SurGE (Survey Generation Evaluation), a new benchmark for scientific survey generation in computer science. SurGE consists of (1) a collection of test instances, each including a topic description, an expert-written survey, and its full set of cited references, and (2) a large-scale academic corpus of over one million papers. In addition, we propose an automated evaluation framework that measures the quality of generated surveys across four dimensions: comprehensiveness, citation accuracy, structural organization, and content quality. Our evaluation of diverse LLM-based methods demonstrates a significant performance gap, revealing that even advanced agentic frameworks struggle with the complexities of survey generation and highlighting the need for future research in this area. We have open-sourced all the code, data, and models at: https://github.com/oneal2000/SurGE

综述生成大模型评估基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。