arXiv:2512.01797cs.AIcs.CL2025-12被引 21

发现语言模型中极少神经元关联幻觉,可预测并控制错误生成。

H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs

  • 仅0.1%的神经元能精准预测幻觉发生,跨场景泛化强。
  • 干预这些神经元可消除过度服从行为,证明其因果作用。
  • 这类神经元源于预训练阶段,提示需从源头优化模型可靠性。

大型语言模型常产生看似合理但事实错误的输出(即幻觉),影响其可信度。现有研究多从训练数据和目标等宏观角度分析幻觉,但对神经元层面的机制仍不清楚。本文从识别、行为影响和起源三方面系统研究了与幻觉相关的神经元(H-Neurons)。在识别上,我们发现极少数神经元(少于总神经元的0.1%)能可靠预测幻觉,且在多种场景下具有强泛化能力;在行为影响上,通过可控干预验证了这些神经元与过度服从行为存在因果关系;在起源上,追踪发现这些神经元源自预训练基础模型,且在后续微调中仍保持预测能力,表明其在预训练阶段已形成。研究将宏观行为模式与微观神经机制联系起来,为构建更可靠的大型语言模型提供了新思路。

原文摘要 · Abstract (English)

Large language models (LLMs) frequently generate hallucinations -- plausible but factually incorrect outputs -- undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored. In this paper, we conduct a systematic investigation into hallucination-associated neurons (H-Neurons) in LLMs from three perspectives: identification, behavioral impact, and origins. Regarding their identification, we demonstrate that a remarkably sparse subset of neurons (less than $0.1\%$ of total neurons) can reliably predict hallucination occurrences, with strong generalization across diverse scenarios. In terms of behavioral impact, controlled interventions reveal that these neurons are causally linked to over-compliance behaviors. Concerning their origins, we trace these neurons back to the pre-trained base models and find that these neurons remain predictive for hallucination detection, indicating they emerge during pre-training. Our findings bridge macroscopic behavioral patterns with microscopic neural mechanisms, offering insights for developing more reliable LLMs.

大模型幻觉神经机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。