arXiv:2511.07689cs.CLcs.AI2025-11ACL被引 2

测试六个事实一致性指标在长文档摘要中的可靠性

Stress Testing Factual Consistency Metrics for Long-Document Summarization

  • 用七种语义不变扰动测试指标稳定性
  • 高信息密度内容下指标可靠性显著下降
  • 适合关注长文档评估的 researchers 参考

评估抽象式文本摘要的事实一致性仍是重大挑战,尤其在长文档场景中,传统指标受限于输入长度和长程依赖。本文系统评估了六个广泛使用的无参考事实性指标在长文档设置下的可靠性。通过七种保持事实性的扰动(改写、简化、同义替换、逻辑等价否定、词汇缩减、压缩、源文插入)测试指标鲁棒性,并分析其对检索上下文和命题信息密度的敏感性。在涵盖科幻、法律和科学领域的三个长文档基准数据集上,结果表明现有短文本指标对语义等价摘要产生不一致评分,且在信息密度高的命题上可靠性下降,其内容与源文本多个部分语义相似。扩大检索上下文虽在某些领域提升稳定性,但无一指标能在长上下文条件下持续保持事实一致性。研究揭示改进方向:多跨度推理、上下文感知校准、基于语义不变变体训练以增强鲁棒性。代码、扰动数据及复现脚本已开源。

原文摘要 · Abstract (English)

Evaluating the factual consistency of abstractive text summarization remains a significant challenge, particularly for long documents, where conventional metrics struggle with input length limitations and long-range dependencies. In this work, we systematically evaluate the reliability of six widely used reference-free factuality metrics, originally proposed for short-form summarization, in the long-document setting. We probe metric robustness through seven factuality-preserving perturbations applied to summaries, namely paraphrasing, simplification, synonym replacement, logically equivalent negations, vocabulary reduction, compression, and source text insertion, and further analyze their sensitivity to retrieval context and claim information density. Across three long-form benchmark datasets spanning science fiction, legal, and scientific domains, our results reveal that existing short-form metrics produce inconsistent scores for semantically equivalent summaries and exhibit declining reliability for information-dense claims whose content is semantically similar to many parts of the source document. While expanding the retrieval context improves stability in some domains, no metric consistently maintains factual alignment under long-context conditions. Finally, our results highlight concrete directions for improving factuality evaluation, including multi-span reasoning, context-aware calibration, and training on meaning-preserving variations to enhance robustness in long-form summarization. We release all code, perturbed data, and scripts required to reproduce our results at https://github.com/zainmujahid/metricEval-longSum.

事实一致性长文档评估指标摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。