发现幻觉神经元无法跨领域通用,提示需按领域定制检测方法。
Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs

- 通过跨领域迁移实验验证幻觉神经元的泛化能力
- 跨域检测性能下降显著(AUROC从0.783降至0.563)
- 适用于需精准识别幻觉的领域特定场景
近期研究识别出一类稀疏的‘幻觉神经元’(H-neurons),其数量不足前馈网络神经元的0.1%,能可靠预测大语言模型是否产生幻觉。这些神经元在通用知识问答任务中被发现,并表现出对新评估样本的泛化能力。我们提出一个自然问题:H-neurons能否在不同知识领域间泛化?通过在6个领域(通用问答、法律、金融、科学、道德推理、代码漏洞)和5个开源权重模型(3B至8B参数)上系统测试跨领域迁移,结果表明它们不具备跨域泛化能力。在某一领域训练的分类器,在跨域测试中仅获得0.563的AUROC,相比域内表现(0.783)下降0.220(p < 0.001),且该趋势在所有模型中一致。结果表明,幻觉并非单一机制或具有通用神经签名,而是依赖于具体知识类型,涉及不同的领域特异性神经元群体。这一发现对神经元级幻觉检测器的部署有直接意义:必须针对每个领域单独校准,而非一次训练全局适用。
原文摘要 · Abstract (English)
Recent work identifies a sparse set of "hallucination neurons" (H-neurons), less than 0.1% of feed-forward network neurons, that reliably predict when large language models will hallucinate. These neurons are identified on general-knowledge question answering and shown to generalize to new evaluation instances. We ask a natural follow-up question: do H-neurons generalize across knowledge domains? Using a systematic cross-domain transfer protocol across 6 domains (general QA, legal, financial, science, moral reasoning, and code vulnerability) and 5 open-weight models (3B to 8B parameters), we find they do not. Classifiers trained on one domain's H-neurons achieve AUROC 0.783 within-domain but only 0.563 when transferred to a different domain (delta = 0.220, p < 0.001), a degradation consistent across all models tested. Our results suggest that hallucination is not a single mechanism with a universal neural signature, but rather involves domain-specific neuron populations that differ depending on the knowledge type being queried. This finding has direct implications for the deployment of neuron-level hallucination detectors, which must be calibrated per domain rather than trained once and applied universally.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。