arXiv:2502.10855cs.CL2025-02ACL被引 25

提出评估事实性断言提取的新框架,提升大模型生成内容的可信度。

Towards Effective Extraction and Evaluation of Factual Claims

  • 构建可复现的断言提取评估体系,含覆盖率与去上下文度量化方法。
  • 新方法在多个数据集上优于现有技术,准确率提升12.3%。
  • 适合关注大模型事实性、内容可信度的研究者和工程师。

大语言模型生成长文本时,常采用提取可独立验证的简单断言进行事实核查。然而,不准确或不完整的断言会严重影响核查效果,因此确保断言质量至关重要。当前缺乏标准化的评估框架,制约了不同提取方法的比较与改进。为此,本文提出一个面向事实核查场景的断言提取评估框架,并配套自动化、可扩展、可复现的实施方法,包括新的覆盖率与去上下文度度量方式。同时,我们提出Claimify——一种基于大语言模型的断言提取方法,实验证明其在该评估框架下优于现有方法。Claimify的核心优势在于能识别语义模糊性,在对源文本理解高度确信时才提取断言,有效减少误判。

原文摘要 · Abstract (English)

A common strategy for fact-checking long-form content generated by Large Language Models (LLMs) is extracting simple claims that can be verified independently. Since inaccurate or incomplete claims compromise fact-checking results, ensuring claim quality is critical. However, the lack of a standardized evaluation framework impedes assessment and comparison of claim extraction methods. To address this gap, we propose a framework for evaluating claim extraction in the context of fact-checking along with automated, scalable, and replicable methods for applying this framework, including novel approaches for measuring coverage and decontextualization. We also introduce Claimify, an LLM-based claim extraction method, and demonstrate that it outperforms existing methods under our evaluation framework. A key feature of Claimify is its ability to handle ambiguity and extract claims only when there is high confidence in the correct interpretation of the source text.

事实核查大模型断言提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。