用问答对细化摘要评估,让内容选择更精准。
QAPyramid: Fine-grained Evaluation of Content Selection for Text Summarization
- 将参考摘要拆分为细粒度问答对进行评估
- 覆盖8.9K个问答级标注,提升评估精度
- 无需专家标注,适合大规模自动化评估
文本摘要的人工评估长期面临挑战。现有的金字塔评估法虽被广泛采用,但子单元定义不系统、粒度粗。本文提出QAPyramid,基于QA-SRL框架将参考摘要分解为更细粒度的问答对,对CNN/DM数据集的参考摘要进行标注,共生成8.9K个问答级标注。实验表明,相比传统金字塔法,QAPyramid在保持高标注者一致性的同时,实现了更系统、更精细的内容选择评估。此外,本文提出的自动化评估指标与QAPyramid的相关性优于现有广泛使用的指标。
原文摘要 · Abstract (English)
How to properly conduct human evaluations for text summarization is a longstanding challenge. The Pyramid human evaluation protocol, which assesses content selection by breaking the reference summary into subunits and verifying their presence in the system summary, has been widely adopted. However, it suffers from a lack of systematicity in the definition and granularity of the sub-units. We address these problems by proposing QAPyramid, which decomposes each reference summary into finer-grained question-answer (QA) pairs according to the QA-SRL framework. We collect QA-SRL annotations for reference summaries from CNN/DM and evaluate 10 summarization systems, resulting in 8.9K QA-level annotations. We show that, compared to Pyramid, QAPyramid provides more systematic and fine-grained content selection evaluation while maintaining high inter-annotator agreement without needing expert annotations. Furthermore, we propose metrics that automate the evaluation pipeline and achieve higher correlations with QAPyramid than other widely adopted metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。