arXiv:2602.08709cs.CL2026-02

提出自动评估观点摘要事实一致性的新方法,提升生成内容可信度。

FactSim: Fact-Checking for Opinion Summarization

  • 通过比对摘要与原文主张的相似性,量化事实一致性。
  • 该方法对否定、改写、扩展等不同表达形式的主张均能正确识别,得分更高。
  • 评分与人工判断高度相关,优于现有主流指标,适合评估生成式摘要质量。

我们探讨了在生成式人工智能文本摘要任务中,尤其是观点摘要领域,需要更全面、精确的评估技术。传统方法依赖自动化指标比较机器生成的摘要与原始意见文本(如产品评论),但随着大语言模型的兴起,这些方法暴露出局限性。本文提出一种全新的全自动评估方法,用于衡量摘要的事实一致性。该方法基于计算摘要中的主张与原始评论中主张之间的相似性,评估生成摘要的覆盖范围与一致性。为此,我们采用一种简单的方法从文本中提取事实主张,再进行对比与汇总,得到可量化的分数。实验表明,该方法对相同主张(无论是否被否定、改写或扩展)均能赋予更高分,且其评分与人工判断具有高相关性,显著优于当前主流评估指标。

原文摘要 · Abstract (English)

We explore the need for more comprehensive and precise evaluation techniques for generative artificial intelligence (GenAI) in text summarization tasks, specifically in the area of opinion summarization. Traditional methods, which leverage automated metrics to compare machine-generated summaries from a collection of opinion pieces, e.g. product reviews, have shown limitations due to the paradigm shift introduced by large language models (LLM). This paper addresses these shortcomings by proposing a novel, fully automated methodology for assessing the factual consistency of such summaries. The method is based on measuring the similarity between the claims in a given summary with those from the original reviews, measuring the coverage and consistency of the generated summary. To do so, we rely on a simple approach to extract factual assessment from texts that we then compare and summarize in a suitable score. We demonstrate that the proposed metric attributes higher scores to similar claims, regardless of whether the claim is negated, paraphrased, or expanded, and that the score has a high correlation to human judgment when compared to state-of-the-art metrics.

观点摘要事实核查评估指标生成式AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。