arXiv:2601.06666cs.CLcs.AI2026-01被引 2

让大模型事实核查更精细可解释,识别错误类型并给出修正建议。

InFi-Check: Interpretable and Fine-Grained Fact-Checking of LLMs

  • 构建带细粒度错误标签的合成数据,支持精准定位问题
  • 在基准测试中超越现有方法,且跨任务泛化能力强
  • 适合需要可信推理和可解释性的AI应用开发者

大型语言模型常产生幻觉,而现有事实核查方法多为二分类,解释性差且忽略错误细节。本文提出InFi-Check框架,首先设计可控数据合成流程,生成包含明确证据、细粒度错误标签、理由与修正的高质量数据。基于此,构建大规模训练数据及人工验证基准InFi-Check-FG,用于细粒度事实核查。在此基础上提出InFi-Checker,可联合输出支持证据、分类错误类型,并生成理由与修正。实验表明,InFi-Checker在InFi-Check-FG上达到领先性能,且在多种下游任务中展现强泛化能力,显著提升事实评估的实用性与可信度。

原文摘要 · Abstract (English)

Large language models (LLMs) often hallucinate, yet most existing fact-checking methods treat factuality evaluation as a binary classification problem, offering limited interpretability and failing to capture fine-grained error types. In this paper, we introduce InFi-Check, a framework for interpretable and fine-grained fact-checking of LLM outputs. Specifically, we first propose a controlled data synthesis pipeline that generates high-quality data featuring explicit evidence, fine-grained error type labels, justifications, and corrections. Based on this, we further construct large-scale training data and a manually verified benchmark InFi-Check-FG for fine-grained fact-checking of LLM outputs. Building on these high-quality training data, we further propose InFi-Checker, which can jointly provide supporting evidence, classify fine-grained error types, and produce justifications along with corrections. Experiments show that InFi-Checker achieves state-of-the-art performance on InFi-Check-FG and strong generalization across various downstream tasks, significantly improving the utility and trustworthiness of factuality evaluation.

事实核查大模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。