arXiv:2605.03971cs.CL2026-05ACL

用逻辑一致性连接回答与自评,提升大模型幻觉检测能力

Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments

论文配图:Logical Consistency as a Bridge: Improving LLM Hallucination Detection via Label Constraint Modeling between Responses and Self-Judgments
图 1 · 摘自论文原文
  • 构建元判断机制,将符号标签映射回特征空间
  • 在4个数据集上优于8种基线方法,显著降低幻觉率
  • 适合关注大模型可信性与自我评估的研究者

大语言模型容易产生事实性幻觉,影响其在实际应用中的可靠性。现有幻觉检测方法主要从微观神经模式中提取不确定性,或通过提示词引导生成宏观自评。但这些方法仅关注幻觉的单一维度,分别处理隐式神经不确定性和显式符号推理,未能利用二者内在关联实现整体感知。本文提出LaaB(Logical Consistency-as-a-Bridge)框架,通过引入“元判断”过程,将符号标签回映至特征空间。基于自评语义中响应与元判断标签相同或相反的逻辑关系,利用双向学习对齐并融合双视角信号,增强幻觉检测能力。在4个公开数据集、4种大模型上,对比8种基线方法的实验表明,LaaB具有显著优势。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are prone to factual hallucinations, risking their reliability in real-world applications. Existing hallucination detectors mainly extract micro-level intrinsic patterns for uncertainty quantification or elicit macro-level self-judgments through verbalized prompts. However, these methods address only a single facet of the hallucination, focusing either on implicit neural uncertainty or explicit symbolic reasoning, thereby treating these inherently coupled behaviors in isolation and failing to exploit their interdependence for a holistic view. In this paper, we propose LaaB (Logical Consistency-as-a-Bridge), a framework that bridges neural features and symbolic judgments for hallucination detection. LaaB introduces a "meta-judgment" process to map symbolic labels back into the feature space. By leveraging the inherent logical bridge where response and meta-judgment labels are either the same or opposite based on the self-judgment's semantics, LaaB aligns and integrates dual-view signals via mutual learning and enhances the hallucination detection. Extensive experiments on 4 public datasets, across 4 LLMs, against 8 baselines demonstrate the superiority of LaaB.

幻觉检测逻辑一致性大模型可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。