提出一种无需参考文本的摘要评估方法,计算快且与人工评价高度相关。
Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics
- 基于无参考文本的评估思路,避免依赖高质量参考摘要
- 在长文档摘要上仍保持与人工评价的高相关性
- 适合资源受限或参考质量差的摘要系统评估场景
自动评估指标常用于替代昂贵的人工标注来评估抽象式摘要系统。理想的指标应具备细粒度、与人工评价高度相关,且不受参考文本质量影响;然而,现有主流摘要评估指标多为依赖参考文本的,而现有无参考指标在长文档摘要上的相关性表现不佳。本文提出一种新的无参考指标,在保持极低计算成本的同时,与人工评价的相关性显著提升。此外,该指标可与参考依赖指标协同使用,增强后者在低质量参考设置下的鲁棒性。
原文摘要 · Abstract (English)
Automatic metrics are used as proxies to evaluate abstractive summarization systems when human annotations are too expensive. To be useful, these metrics should be fine-grained, show a high correlation with human annotations, and ideally be independent of reference quality; however, most standard evaluation metrics for summarization are reference-based, and existing reference-free metrics correlate poorly with relevance, especially on summaries of longer documents. In this paper, we introduce a reference-free metric that correlates well with human evaluated relevance, while being very cheap to compute. We show that this metric can also be used alongside reference-based metrics to improve their robustness in low quality reference settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。