语言模型幻觉源于不确定信息未融入输出生成,而非检测失败。
The Phenomenology of Hallucinations
- 不确定输入被准确识别,占据高维空间且维度是真实输入的2-3倍。
- 不确定信号在输出层弱耦合,几何放大却功能沉默,无法引发拒绝输出。
- 因果干预证明:直接连接不确定性到输出层可恢复拒绝机制,验证核心机制。
我们发现语言模型产生幻觉并非因无法检测不确定性,而是由于未能将该不确定性整合进输出生成过程。在多种架构中,不确定输入均能被可靠识别,占据高维空间,其内在维度达真实输入的2至3倍。然而,这一内部信号与输出层耦合极弱:不确定性迁移到低敏感性子空间,虽被几何放大却功能上处于沉默状态。拓扑分析显示,不确定性表征呈现碎片化,而非汇聚为统一的拒答状态;梯度与费雪探测器揭示沿不确定性方向敏感性持续坍缩。由于交叉熵训练无拒答吸引子,且对自信预测一视同仁地奖励,关联机制会不断放大这些断裂的激活,直至残余耦合迫使模型生成确定性输出。因果干预证实此理论:当不确定性被直接连接至逻辑输出时,模型可恢复拒绝行为。
原文摘要 · Abstract (English)
We show that language models hallucinate not because they fail to detect uncertainty, but because of a failure to integrate it into output generation. Across architectures, uncertain inputs are reliably identified, occupying high-dimensional regions with 2-3$\times$ the intrinsic dimensionality of factual inputs. However, this internal signal is weakly coupled to the output layer: uncertainty migrates into low-sensitivity subspaces, becoming geometrically amplified yet functionally silent. Topological analysis shows that uncertainty representations fragment rather than converging to a unified abstention state, while gradient and Fisher probes reveal collapsing sensitivity along the uncertainty direction. Because cross-entropy training provides no attractor for abstention and uniformly rewards confident prediction, associative mechanisms amplify these fractured activations until residual coupling forces a committed output despite internal detection. Causal interventions confirm this account by restoring refusal when uncertainty is directly connected to logits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。