arXiv:2606.08705cs.CL2026-06

探究大模型幻觉与知识冲突的关联,发现二者相关但不等同。

Analyzing the Correlation Between Hallucinations and Knowledge Conflicts in Large Language Models

  • 通过探测隐藏层、注意力和MLP层激活值分析内部表征
  • 发现幻觉模式无法完全由知识冲突解释,但存在部分相关性
  • 验证探测技术在多语言、多层结构中的有效性,助力模型可解释性

幻觉——即事实错误或不可验证的输出——仍是大语言模型(LLMs)在知识密集型任务中的主要挑战之一。一种假设认为,其源于固定且过时训练数据引发的内部知识冲突。本文研究了与知识冲突相关的内部表征是否与幻觉行为存在关联。基于两项先前工作启发的探测方法,我们分析了预定义任务下隐层、注意力层、MLP层及输出逻辑值的激活情况。实验对LLaMA-3-8B在幻觉检测基准上进行,对Falcon-7B在知识冲突数据集上进行。结果表明,尽管概念相关,幻觉激活模式不能被知识冲突表征完全还原或解释。然而,探测技术在多种语言和激活类型中表现稳健,支持其在提升大模型可解释性方面的价值。本研究深化了对大模型幻觉机制的理解,并强调了对其内部行为进行细粒度分析的重要性。

原文摘要 · Abstract (English)

Hallucinations -- factually incorrect or unverifiable outputs -- remain one of the most challenging limitations of Large Language Models (LLMs), especially in knowledge-intensive tasks. One proposed explanation is internal knowledge conflicts arising from fixed, outdated training data. This paper investigates whether internal representations linked to knowledge conflicts correlate with hallucination behaviors in LLMs. Using probing techniques inspired by two prior works, we analyzed activations from hidden, attention, and MLP layers, as well as output logits, across predefined tasks. We probed LLaMA-3-8B on hallucination detection benchmarks and Falcon-7B on a knowledge conflict dataset. Our findings show that, although conceptually related, hallucination activation patterns cannot be fully reduced to or explained by knowledge conflict representations. Nonetheless, probing proves a robust tool across multiple languages and activation types, supporting its role in improving LLM interpretability. This work advances the broader understanding of hallucinations in LLMs and underscores the value of fine-grained analysis of their internal behavior.

大模型幻觉可解释性知识冲突

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。