arXiv:2512.20182cs.CLcs.AI2025-12ACL被引 5

FaithLens可检测并解释大模型幻觉,提升可信度。

FaithLens: Detecting and Explaining Faithfulness Hallucination

论文配图:FaithLens: Detecting and Explaining Faithfulness Hallucination
图 1 · 摘自论文原文
  • 用大模型生成带解释的训练数据,再通过筛选确保质量
  • 80亿参数模型在12项任务上超越GPT-5.2和o3
  • 兼具高可信、高效、可解释性,适合实际部署

识别大语言模型输出中的忠实性幻觉对检索增强生成、摘要等实际应用至关重要。本文提出FaithLens,一种低成本且高效的忠实性幻觉检测模型,可同时提供二分类预测与对应解释,以提升可信度。我们首先利用先进大模型合成带解释的训练数据,并采用严格的数据过滤策略,确保标签正确性、解释质量和数据多样性。随后,在高质量数据上微调模型作为冷启动,并进一步通过基于规则的强化学习优化,奖励预测准确性和解释质量。在12个多样化任务上的实验表明,80亿参数的FaithLens性能优于GPT-5.2和o3等先进模型。同时,FaithLens能生成高质量解释,实现可信度、效率与有效性的独特平衡。

原文摘要 · Abstract (English)

Recognizing whether outputs from large language models (LLMs) contain faithfulness hallucination is crucial for real-world applications, e.g., retrieval-augmented generation and summarization. In this paper, we introduce FaithLens, a cost-efficient and effective faithfulness hallucination detection model that can jointly provide binary predictions and corresponding explanations to improve trustworthiness. To achieve this, we first synthesize training data with explanations via advanced LLMs and apply a well-defined data filtering strategy to ensure label correctness, explanation quality, and data diversity. Subsequently, we fine-tune the model on these well-curated training data as a cold start and further optimize it with rule-based reinforcement learning, using rewards for both prediction correctness and explanation quality. Results on 12 diverse tasks show that the 8B-parameter FaithLens outperforms advanced models such as GPT-5.2 and o3. Also, FaithLens can produce high-quality explanations, delivering a distinctive balance of trustworthiness, efficiency, and effectiveness.

大模型幻觉检测可解释性可信生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。