arXiv:2607.00158cs.CL2026-07

医学大模型幻觉可被检测但难控制,内部信号分布广泛难以精准干预。

Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination

论文配图:Readable but Not Controllable: Neuron-Level Evidence for Medical LLM Hallucination
图 1 · 摘自论文原文
  • 通过精心设计的探针检测幻觉,准确率AUROC达0.77~0.86。
  • 随机选择数百个神经元即可恢复大部分检测信号,表明信息分布冗余。
  • 检测易、控制难,揭示了内部表征可见性与可操控性间的根本鸿沟。

在多个医学问答数据集上,使用四个开源模型验证发现:通过精心设计的探针可可靠检测医疗大模型的幻觉,其AUROC分数介于0.77至0.86之间。该检测信号呈分布式且冗余,小规模随机选取的神经元子集已能恢复接近完整的检测性能,低维随机投影亦保留主要识别能力。进一步测试显示,在16组模型-数据集组合中,幻觉的可解码性与可控制性之间存在显著差距——能够被检测的内部表征,却无法用于有效控制或纠正幻觉。这表明医学幻觉虽在激活值中明显可见,但仅靠调控相关神经元难以实现可靠修正。研究提示,幻觉缓解不能仅依赖定位特定神经元,背后反映的是表征揭示能力与实际干预能力之间的深层分离。

原文摘要 · Abstract (English)

Hallucination remains one of the central obstacles to deploying medical LLMs. Yet, even when hallucination can be detected, it is still unclear whether the internal representations associated with it can be used for control rather than detection alone. Using four open-source models across a suite of medical question-answering datasets, we show that a simple, carefully conditioned probe can reliably detect hallucination, with AUROC scores between 0.77 and 0.86 in our case. We further show that this signal is distributed and redundant rather than narrowly localized. Systematically selected neurons outperform random neurons only at very small subset sizes, whereas random subsets of a few hundred neurons recover nearly the full signal, and low-dimensional random projections preserve most of the detection performance. Beyond detection, we test whether this representation is causally actionable. Across 16 model--dataset combinations, our results reveal a sharp gap between decodability and controllability. The same internal structure that makes hallucination easy to detect does not translate into reliable neuron-level control. These findings show that medical hallucination seems to be readily visible in internal activations, but not easily corrected by steering the neurons most associated with it. More broadly, our results suggest that hallucination mitigation is not simply a matter of identifying the right neurons, and point to a deeper separation between what representations reveal and what they allow us to change.

医学LLM幻觉检测神经元控制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。