用事件重合度评估新闻摘要质量,更贴近人类理解。
Event-based evaluation of abstractive news summarization
- 以事件为单位计算生成摘要与原文、参考摘要的重合度
- 在挪威语数据集上验证,显著提升评估可信度
- 适合关注摘要信息完整性与事件覆盖的研究者
新闻摘要通过浓缩形式包含文章的核心信息。当前自动摘要的评估高度依赖人工撰写的参考摘要,通过计算重叠单元或相似度得分进行判断。新闻报道的是事件,理想情况下摘要也应如此呈现。本文提出一种新方法:通过计算生成摘要、参考摘要与原始新闻文章之间的事件重合度来评估摘要质量。我们在一个标注丰富的挪威语数据集上进行了实验,该数据集包含事件标注和专家撰写的摘要。结果表明,该方法能更深入揭示摘要中包含的事件信息,提供比传统方法更可靠的评估视角。
原文摘要 · Abstract (English)
An abstractive summary of a news article contains its most important information in a condensed version. The evaluation of automatically generated summaries by generative language models relies heavily on human-authored summaries as gold references, by calculating overlapping units or similarity scores. News articles report events, and ideally so should the summaries. In this work, we propose to evaluate the quality of abstractive summaries by calculating overlapping events between generated summaries, reference summaries, and the original news articles. We experiment on a richly annotated Norwegian dataset comprising both events annotations and summaries authored by expert human annotators. Our approach provides more insight into the event information contained in the summaries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。