用话语分析提升长文摘要事实一致性检测能力
Discourse-Driven Evaluation: Unveiling Factual Inconsistency in Long Document Summarization
- 将长文本按话语特征切分为片段,增强结构化理解
- 在多个基准上优于基线模型,提升事实一致性识别率
- 适合关注长文档摘要质量评估的研究者和开发者
长文档摘要的事实一致性检测仍具挑战,源于原文结构复杂与摘要长度长。本文研究事实不一致错误,并将其与话语分析关联。发现错误更常出现在复杂句中,且与多种话语特征相关。提出一种框架,将长文本分解为受话语启发的片段,并利用话语信息整合自然语言推理模型预测的句子级评分。该方法在多个涵盖丰富文本领域的评估基准上,优于不同基线模型,凸显在长文档摘要事实一致性评分中引入话语特征的重要性。
原文摘要 · Abstract (English)
Detecting factual inconsistency for long document summarization remains challenging, given the complex structure of the source article and long summary length. In this work, we study factual inconsistency errors and connect them with a line of discourse analysis. We find that errors are more common in complex sentences and are associated with several discourse features. We propose a framework that decomposes long texts into discourse-inspired chunks and utilizes discourse information to better aggregate sentence-level scores predicted by natural language inference models. Our approach shows improved performance on top of different model baselines over several evaluation benchmarks, covering rich domains of texts, focusing on long document summarization. This underscores the significance of incorporating discourse features in developing models for scoring summaries for long document factual inconsistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。