arXiv:2409.19507cs.CL2024-09EMNLP被引 2

现有摘要评估指标的元评估存在数据单一、忽视忠实性的问题,亟需更全面的评测基准。

A Critical Look at Meta-evaluating Summarisation Evaluation Metrics

  • 基于新闻摘要数据集进行元评估,覆盖范围有限。
  • 研究重点转向生成摘要的忠实性评估,但缺乏多样性。
  • 建议构建多样化基准,提升评估指标的鲁棒性与通用性。

有效的摘要评估指标能帮助研究人员和实践者高效比较不同摘要系统。评估自动评估指标的有效性(即元评估)是一个关键研究问题。本文回顾了近期摘要评估指标的元评估实践,发现:(1) 评估指标主要在新闻摘要数据集样本上进行元评估;(2) 研究焦点明显转向生成摘要的忠实性评估。我们主张应尽快建立更丰富的评测基准,以促进更稳健评估指标的发展,并分析现有评估指标的泛化能力。此外,呼吁开展关注用户需求的质量维度研究,考虑摘要的传播目标及其在工作流中的角色。

原文摘要 · Abstract (English)

Effective summarisation evaluation metrics enable researchers and practitioners to compare different summarisation systems efficiently. Estimating the effectiveness of an automatic evaluation metric, termed meta-evaluation, is a critically important research question. In this position paper, we review recent meta-evaluation practices for summarisation evaluation metrics and find that (1) evaluation metrics are primarily meta-evaluated on datasets consisting of examples from news summarisation datasets, and (2) there has been a noticeable shift in research focus towards evaluating the faithfulness of generated summaries. We argue that the time is ripe to build more diverse benchmarks that enable the development of more robust evaluation metrics and analyze the generalization ability of existing evaluation metrics. In addition, we call for research focusing on user-centric quality dimensions that consider the generated summary's communicative goal and the role of summarisation in the workflow.

摘要评估元评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。