arXiv:2411.05980cs.CLcs.AI2024-11ACL被引 11

提出细粒度事实验证基准,精准识别大模型生成内容中的错误

FactLens: Benchmarking Fine-Grained Fact Verification

  • 将复杂陈述拆解为子陈述逐一验证,提升准确性
  • 人工构建高质量标注数据,确保验证结果可信
  • 自动化评估器与人工判断高度一致,适合模型评测

大型语言模型在生成和理解语言方面表现出色,但其幻觉问题导致事实性错误仍是一大瓶颈。传统验证方法通常对复杂陈述给出单一真假标签,可能掩盖细微错误。本文倡导细粒度验证:将复杂陈述分解为更小的子陈述进行独立验证,从而更精确地定位错误、增强透明度并减少证据检索歧义。然而,子陈述的生成面临保持上下文连贯性和语义等价性的挑战。为此,我们提出了FactLens——一个用于评估细粒度事实验证的基准,包含评估指标和自动化子陈述质量评估工具。数据集经过人工精心标注,确保高质量真实标签。实验表明,自动化评估器与人工判断高度一致,并讨论了子陈述特征对整体验证性能的影响。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have shown impressive capability in language generation and understanding, but their tendency to hallucinate and produce factually incorrect information remains a key limitation. To verify LLM-generated contents and claims from other sources, traditional verification approaches often rely on holistic models that assign a single factuality label to complex claims, potentially obscuring nuanced errors. In this paper, we advocate for a shift towards fine-grained verification, where complex claims are broken down into smaller sub-claims for individual verification, allowing for more precise identification of inaccuracies, improved transparency, and reduced ambiguity in evidence retrieval. However, generating sub-claims poses challenges, such as maintaining context and ensuring semantic equivalence with respect to the original claim. We introduce FactLens, a benchmark for evaluating fine-grained fact verification, with metrics and automated evaluators of sub-claim quality. The benchmark data is manually curated to ensure high-quality ground truth. Our results show alignment between automated FactLens evaluators and human judgments, and we discuss the impact of sub-claim characteristics on the overall verification performance.

事实验证大模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。