构建多维度评估框架,提升自动综述生成系统评测可靠性
SGSimEval: A Comprehensive Multifaceted and Similarity-Enhanced Benchmark for Automatic Survey Generation Systems
- 融合大纲、内容、参考文献的多维度评估方法
- 结合人工偏好与量化指标,避免依赖LLM评分
- 适合研究自动综述生成与评测的学者使用
自动综述生成(ASG)因大语言模型(LLMs)的发展而受到关注,传统需大量人力的工作如今可通过检索增强生成(RAG)和多智能体系统(MASs)实现。然而现有评估方法存在度量偏差、缺乏人类偏好、过度依赖LLM作为评判者等问题。为此,我们提出SGSimEval,一个综合性多维且基于相似性的评估基准,通过整合大纲、内容、参考文献的评估,并结合基于LLM的评分与定量指标,构建多角度评价体系。同时引入强调内在质量与人类相似性的用户偏好度量。大量实验表明,当前ASG系统在大纲生成上已接近人类水平,但在内容和参考文献生成方面仍有显著提升空间,且本评估方法与人工评估高度一致。
原文摘要 · Abstract (English)
The growing interest in automatic survey generation (ASG), a task that traditionally required considerable time and effort, has been spurred by recent advances in large language models (LLMs). With advancements in retrieval-augmented generation (RAG) and the rising popularity of multi-agent systems (MASs), synthesizing academic surveys using LLMs has become a viable approach, thereby elevating the need for robust evaluation methods in this domain. However, existing evaluation methods suffer from several limitations, including biased metrics, a lack of human preference, and an over-reliance on LLMs-as-judges. To address these challenges, we propose SGSimEval, a comprehensive benchmark for Survey Generation with Similarity-Enhanced Evaluation that evaluates automatic survey generation systems by integrating assessments of the outline, content, and references, and also combines LLM-based scoring with quantitative metrics to provide a multifaceted evaluation framework. In SGSimEval, we also introduce human preference metrics that emphasize both inherent quality and similarity to humans. Extensive experiments reveal that current ASG systems demonstrate human-comparable superiority in outline generation, while showing significant room for improvement in content and reference generation, and our evaluation metrics maintain strong consistency with human assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。