arXiv:2602.07253cs.AIcs.CL2026-02被引 2

将幻觉检测重新定义为分布外检测,提升大模型推理任务的可靠性。

From Out-of-Distribution Detection to Hallucination Detection: A Geometric View

  • 把语言模型的下一步词预测看作分类问题,用分布外检测方法识别幻觉。
  • 无需训练、单样本即可检测,推理任务中幻觉检测准确率显著提升。
  • 适合关注大模型安全与可靠性的研究者,尤其在复杂推理场景下。

大语言模型中的幻觉检测是一个关键且开放的问题,对安全性与可靠性具有重要影响。尽管现有幻觉检测方法在问答任务中表现良好,但在需要推理的任务上效果仍不理想。本文从分布外(OOD)检测的视角重新审视幻觉检测,这一问题在计算机视觉等领域已有深入研究。通过将语言模型的下一步词预测视为分类任务,可应用OOD检测技术,只需针对大模型的结构差异进行适当调整。实验表明,基于OOD的方法能实现无需训练、仅需单样本的检测器,在推理任务中表现出优异的幻觉检测准确性。整体而言,将幻觉检测重构为分布外检测,为提升语言模型安全性提供了一条有前景且可扩展的路径。

原文摘要 · Abstract (English)

Detecting hallucinations in large language models is a critical open problem with significant implications for safety and reliability. While existing hallucination detection methods achieve strong performance in question-answering tasks, they remain less effective on tasks requiring reasoning. In this work, we revisit hallucination detection through the lens of out-of-distribution (OOD) detection, a well-studied problem in areas like computer vision. Treating next-token prediction in language models as a classification task allows us to apply OOD techniques, provided appropriate modifications are made to account for the structural differences in large language models. We show that OOD-based approaches yield training-free, single-sample-based detectors, achieving strong accuracy in hallucination detection for reasoning tasks. Overall, our work suggests that reframing hallucination detection as OOD detection provides a promising and scalable pathway toward language model safety.

幻觉检测分布外检测大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。