用大模型实现细粒度摘要评估,提升客观性和可解释性。
Evaluate Summarization in Fine-Granularity: Auto Evaluation with LLM
- 通过多粒度维度评估摘要的完整性、正确性等四方面表现。
- 相比传统方法,与人工评分相关性显著提升。
- 适合需要精准评估摘要质量的研究者与开发者。
随着信息爆炸和高效信息消费的需求,摘要任务日益重要。然而,对长篇且内容丰富的非结构化文本进行准确、客观的摘要评估仍面临巨大挑战。现有方法如ROUGE(Lin, 2004)和嵌入相似性,其得分与人类判断相关性低,且难以理解,无法真实反映摘要质量。虽然大语言模型(LLM)可模拟人类主观评价,但主观评分难以解释且易受模型和提示风格影响。本文提出一种新型评估方法与工具SumAutoEval,可在不同粒度层级上评估摘要在完整性、正确性、一致性与可读性四个关键维度的表现,提供客观评分。实验证明,该方法显著提升了对摘要质量的理解,并与人工评分具有更高相关性。
原文摘要 · Abstract (English)
Due to the exponential growth of information and the need for efficient information consumption the task of summarization has gained paramount importance. Evaluating summarization accurately and objectively presents significant challenges, particularly when dealing with long and unstructured texts rich in content. Existing methods, such as ROUGE (Lin, 2004) and embedding similarities, often yield scores that have low correlation with human judgements and are also not intuitively understandable, making it difficult to gauge the true quality of the summaries. LLMs can mimic human in giving subjective reviews but subjective scores are hard to interpret and justify. They can be easily manipulated by altering the models and the tones of the prompts. In this paper, we introduce a novel evaluation methodology and tooling designed to address these challenges, providing a more comprehensive, accurate and interpretable assessment of summarization outputs. Our method (SumAutoEval) proposes and evaluates metrics at varying granularity levels, giving objective scores on 4 key dimensions such as completeness, correctness, Alignment and readability. We empirically demonstrate, that SumAutoEval enhances the understanding of output quality with better human correlation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。