arXiv:2410.10867cs.CLcs.AI2024-10EMNLP被引 7

提出一种无需参考文本的摘要评估方法,计算快且与人工评价高度相关。

Mitigating the Impact of Reference Quality on Evaluation of Summarization Systems with Reference-Free Metrics

  • 基于无参考文本的评估思路,避免依赖高质量参考摘要
  • 在长文档摘要上仍保持与人工评价的高相关性
  • 适合资源受限或参考质量差的摘要系统评估场景

自动评估指标常用于替代昂贵的人工标注来评估抽象式摘要系统。理想的指标应具备细粒度、与人工评价高度相关,且不受参考文本质量影响;然而,现有主流摘要评估指标多为依赖参考文本的,而现有无参考指标在长文档摘要上的相关性表现不佳。本文提出一种新的无参考指标,在保持极低计算成本的同时,与人工评价的相关性显著提升。此外,该指标可与参考依赖指标协同使用,增强后者在低质量参考设置下的鲁棒性。

原文摘要 · Abstract (English)

Automatic metrics are used as proxies to evaluate abstractive summarization systems when human annotations are too expensive. To be useful, these metrics should be fine-grained, show a high correlation with human annotations, and ideally be independent of reference quality; however, most standard evaluation metrics for summarization are reference-based, and existing reference-free metrics correlate poorly with relevance, especially on summaries of longer documents. In this paper, we introduce a reference-free metric that correlates well with human evaluated relevance, while being very cheap to compute. We show that this metric can also be used alongside reference-based metrics to improve their robustness in low quality reference settings.

摘要评估无参考自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。