arXiv:2512.02763cs.CLcs.AI2025-12被引 2

构建全面评估大模型生成综述质量的基准,提升自动化文献调研可信度。

SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys

  • 从整体质量、结构连贯性、参考准确性三方面评估自动生成综述。
  • 7个学科测试显示专业综述系统表现远超通用写作模型。
  • 引入人工参考增强评估与人类判断一致性,适合研究自动化文献生成者。

基于大模型的自动综述系统通过整合检索、组织与内容合成,实现端到端的网页信息获取。尽管近期研究聚焦于新生成流程,但如何评估此类复杂系统仍是重大挑战。为此,我们提出SurveyEval,一个涵盖三个维度的综合性评估基准:整体质量、结构连贯性与参考准确性。评估扩展至7个学科,并在LLM-as-a-Judge框架中引入人工参考,以增强评估与人类判断的一致性。实验结果表明,通用长文本或论文写作系统生成的综述质量较低,而专用综述生成系统则显著提升产出质量。我们期望SurveyEval成为跨学科、多标准评估自动综述系统的可扩展测试平台。

原文摘要 · Abstract (English)

LLM-based automatic survey systems are transforming how users acquire information from the web by integrating retrieval, organization, and content synthesis into end-to-end generation pipelines. While recent works focus on developing new generation pipelines, how to evaluate such complex systems remains a significant challenge. To this end, we introduce SurveyEval, a comprehensive benchmark that evaluates automatically generated surveys across three dimensions: overall quality, outline coherence, and reference accuracy. We extend the evaluation across 7 subjects and augment the LLM-as-a-Judge framework with human references to strengthen evaluation-human alignment. Evaluation results show that while general long-text or paper-writing systems tend to produce lower-quality surveys, specialized survey-generation systems are able to deliver substantially higher-quality results. We envision SurveyEval as a scalable testbed to understand and improve automatic survey systems across diverse subjects and evaluation criteria.

大模型评估综述生成自动化文献调研基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。