构建全面评估大模型生成综述质量的基准,提升自动化文献调研可信度。
SurveyEval: Towards Comprehensive Evaluation of LLM-Generated Academic Surveys
- 从整体质量、结构连贯性、参考准确性三方面评估自动生成综述。
- 7个学科测试显示专业综述系统表现远超通用写作模型。
- 引入人工参考增强评估与人类判断一致性,适合研究自动化文献生成者。
基于大模型的自动综述系统通过整合检索、组织与内容合成,实现端到端的网页信息获取。尽管近期研究聚焦于新生成流程,但如何评估此类复杂系统仍是重大挑战。为此,我们提出SurveyEval,一个涵盖三个维度的综合性评估基准:整体质量、结构连贯性与参考准确性。评估扩展至7个学科,并在LLM-as-a-Judge框架中引入人工参考,以增强评估与人类判断的一致性。实验结果表明,通用长文本或论文写作系统生成的综述质量较低,而专用综述生成系统则显著提升产出质量。我们期望SurveyEval成为跨学科、多标准评估自动综述系统的可扩展测试平台。
原文摘要 · Abstract (English)
LLM-based automatic survey systems are transforming how users acquire information from the web by integrating retrieval, organization, and content synthesis into end-to-end generation pipelines. While recent works focus on developing new generation pipelines, how to evaluate such complex systems remains a significant challenge. To this end, we introduce SurveyEval, a comprehensive benchmark that evaluates automatically generated surveys across three dimensions: overall quality, outline coherence, and reference accuracy. We extend the evaluation across 7 subjects and augment the LLM-as-a-Judge framework with human references to strengthen evaluation-human alignment. Evaluation results show that while general long-text or paper-writing systems tend to produce lower-quality surveys, specialized survey-generation systems are able to deliver substantially higher-quality results. We envision SurveyEval as a scalable testbed to understand and improve automatic survey systems across diverse subjects and evaluation criteria.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。