NOIR自动评估摘要质量,兼顾信息保留与长度压缩。
An Automated Length-Aware Quality Metric for Summarization
- 基于语义相似度与长度压缩率,量化摘要的召回-压缩平衡。
- 与人工评价高度相关,无需参考摘要即可评估。
- 适合优化摘要模型、提示词和合成摘要的质量。
本文提出一种名为NOIR(NOrmed Index of Retention)的自动化定量评估指标,用于衡量任意文本摘要的质量。该指标同时考虑语义信息保留程度与摘要长度压缩效果,从而反映摘要生成中召回率与压缩率之间的权衡能力,这是摘要任务的核心技能。实验表明,NOIR能有效捕捉摘要模型在词元长度与语义保留间的权衡关系,并与人类对摘要质量的感知高度相关。该方法采用语言模型嵌入计算语义相似度,无需依赖耗时的人工参考摘要,可广泛应用于各类摘要任务,为评估和改进摘要算法、提示词设计及合成摘要提供自动化工具。
原文摘要 · Abstract (English)
This paper proposes NOrmed Index of Retention (NOIR), a quantitative objective metric for evaluating summarization quality of arbitrary texts that relies on both the retention of semantic meaning and the summary length compression. This gives a measure of how well the recall-compression tradeoff is managed, the most important skill in summarization. Experiments demonstrate that NOIR effectively captures the token-length / semantic retention tradeoff of a summarizer and correlates to human perception of sumarization quality. Using a language model-embedding to measure semantic similarity, it provides an automated alternative for assessing summarization quality without relying on time-consuming human-generated reference summaries. The proposed metric can be applied to various summarization tasks, offering an automated tool for evaluating and improving summarization algorithms, summarization prompts, and synthetically-generated summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。