用高质量数据集和评估框架提升AI生成综述的可信度
SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
- 构建包含4200篇人工撰写综述的大规模数据集
- 引入质量感知指标,提升引用文献筛选准确性
- 验证了全自动生成仍存在引用质量低的问题
自动综述生成已成为科学文档处理中的关键任务。尽管大语言模型在生成综述方面展现出潜力,但缺乏标准化评估数据集严重制约了其性能的严格评估。本文提出SurveyGen,一个涵盖4200余篇跨领域人工撰写综述的大型数据集,包含242,143条引文及详尽的质量相关元数据(涵盖综述与被引论文)。基于此资源,我们构建了QUAL-SG——一种质量感知的综述生成框架,通过在标准检索增强生成(RAG)流程中引入质量感知指标,优化文献检索与选择,以获取更高质量的源论文。利用该数据集与框架,我们系统评估了前沿大模型在不同人机协作程度下的表现:从完全自动到人机协同。实验结果与人工评估显示,半自动流程可达到部分可比效果,但全自动生成仍存在引用质量低、批判性分析不足等问题。
原文摘要 · Abstract (English)
Automatic survey generation has emerged as a key task in scientific document processing. While large language models (LLMs) have shown promise in generating survey texts, the lack of standardized evaluation datasets critically hampers rigorous assessment of their performance against human-written surveys. In this work, we present SurveyGen, a large-scale dataset comprising over 4,200 human-written surveys across diverse scientific domains, along with 242,143 cited references and extensive quality-related metadata for both the surveys and the cited papers. Leveraging this resource, we build QUAL-SG, a novel quality-aware framework for survey generation that enhances the standard Retrieval-Augmented Generation (RAG) pipeline by incorporating quality-aware indicators into literature retrieval to assess and select higher-quality source papers. Using this dataset and framework, we systematically evaluate state-of-the-art LLMs under varying levels of human involvement - from fully automatic generation to human-guided writing. Experimental results and human evaluations show that while semi-automatic pipelines can achieve partially competitive outcomes, fully automatic survey generation still suffers from low citation quality and limited critical analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。