发现大模型虚构引用主要源于特定神经元,可精准干预修复。
Where Fake Citations Are Made: Tracing Field-Level Hallucination to Specific Neurons in LLMs

- 通过分析9个模型10.8万条引用,定位到作者名错误率最高。
- 识别出仅占极少数的场域特异性幻觉神经元(FH-neurons)。
- 抑制这些神经元能显著降低虚构引用,适合模型可解释性研究者。
大型语言模型频繁生成看似合理实则虚构的引用,且常表现出高置信度。我们对9个模型和108,000条生成引用进行了研究,发现作者名错误率远高于其他字段,且引用格式无显著影响;基于推理的蒸馏会降低召回率。在不同字段间训练的探测器转移性能接近随机水平,表明幻觉信号不具备跨字段泛化能力。在此基础上,我们对Qwen2.5-32B-Instruct的神经元级CETT值施加弹性网络正则化与稳定性选择,识别出一组稀疏的场域特异性幻觉神经元(FH-neurons)。因果干预验证其作用:增强这些神经元会提升幻觉,抑制则改善各领域表现,尤其在部分领域提升显著。结果表明,仅利用模型内部信号即可实现轻量级的引用幻觉检测与缓解。
原文摘要 · Abstract (English)
LLMs frequently generate fictitious yet convincing citations, often expressing high confidence even when the underlying reference is wrong. We study this failure across 9 models and 108{,}000 generated references, and find that author names fail far more often than other fields across all models and settings. Citation style has no measurable effect, while reasoning-oriented distillation degrades recall. Probes trained on one field transfer at near-chance levels to the others, suggesting that hallucination signals do not generalize across fields. Building on this finding, we apply elastic-net regularization with stability selection to neuron-level CETT values of Qwen2.5-32B-Instruct and identify a sparse set of field-specific hallucination neurons (FH-neurons). Causal intervention further confirms their role: amplifying these neurons increases hallucination, while suppressing them improves performance across fields, with larger gains in some fields. These results suggest a lightweight approach to detecting and mitigating citation hallucination using internal model signals alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。