arXiv:2509.04483cs.CLcs.AI2025-09被引 3

提出新评估指标,让大模型生成的查证结论更可信

DecMetrics: Structured Claim Decomposition Scoring for Factually Consistent LLM Outputs

  • 设计三个自动评估指标,分别衡量分解结果的完整度、正确性和语义清晰度
  • 用这些指标作为奖励函数优化模型,提升分解结果质量
  • 适合关注事实核查系统可靠性的研究者和开发者

主张分解在事实核查中至关重要,能将复杂声明拆分为简单原子成分并识别错误部分。尽管如此,现有研究多聚焦于生成方法,对分解结果质量的评估仍显不足。为此,本文提出DecMetrics,包含三个新指标:COMPLETENESS(完整度)、CORRECTNESS(正确性)和SEMANTIC ENTROPY(语义熵),用于自动评估分解模型生成的声明质量。基于这些指标,我们构建了一个轻量级声明分解模型,并通过将其作为奖励函数进行优化,以提升性能。本方法旨在建立声明分解的自动化评估基准,增强事实核查系统的可靠性与有效性。

原文摘要 · Abstract (English)

Claim decomposition plays a crucial role in the fact-checking process by breaking down complex claims into simpler atomic components and identifying their unfactual elements. Despite its importance, current research primarily focuses on generative methods for decomposition, with insufficient emphasis on evaluating the quality of these decomposed atomic claims. To bridge this gap, we introduce \textbf{DecMetrics}, which comprises three new metrics: \texttt{COMPLETENESS}, \texttt{CORRECTNESS}, and \texttt{SEMANTIC ENTROPY}, designed to automatically assess the quality of claims produced by decomposition models. Utilizing these metrics, we develop a lightweight claim decomposition model, optimizing its performance through the integration of these metrics as a reward function. Through automatic evaluation, our approach aims to set a benchmark for claim decomposition, enhancing both the reliability and effectiveness of fact-checking systems.

事实核查大模型评估自然语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。