arXiv:2509.06938cs.LGcs.AI2025-09NeurIPS被引 6

揭示大模型幻觉产生的根源:噪声输入会激活无意义的语义概念。

From Noise to Narrative: Tracing the Origins of Hallucinations in Transformers

  • 用稀疏自编码器捕捉模型内部概念,追踪输入不确定性如何引发幻觉
  • 纯噪声输入下仍触发大量稳定且有意义的语义概念,证实幻觉可预测
  • 为安全对齐、风险量化和对抗防御提供新思路,适合关注AI可信性的研究者

随着生成式AI在科学、商业和政府领域的普及,对其故障模式的深入理解变得尤为迫切。变压器模型容易产生幻觉,这种行为波动阻碍了高风险领域对AI解决方案的信任与采纳。本文通过实验控制输入空间中的不确定性,利用稀疏自编码器捕获的语义概念表征,系统揭示了预训练变压器模型中幻觉的成因。实验表明,当输入信息逐渐无结构化时,模型使用的语义概念数量增加。在输入不确定性加剧的情况下,模型倾向于激活与输入无关但连贯的语义特征,导致输出幻觉。极端情况下,对于纯噪声输入,我们识别出预训练变压器模型中间层激活中广泛存在的、可复现且有意义的概念,其功能完整性通过定向调控得到验证。此外,我们证明可通过分析变压器层激活中的概念模式可靠预测模型输出的幻觉。这些关于变压器内部处理机制的洞见,对对齐人工智能与人类价值观、保障AI安全、防范潜在对抗攻击以及实现模型幻觉风险的自动化量化具有直接意义。

原文摘要 · Abstract (English)

As generative AI systems become competent and democratized in science, business, and government, deeper insight into their failure modes now poses an acute need. The occasional volatility in their behavior, such as the propensity of transformer models to hallucinate, impedes trust and adoption of emerging AI solutions in high-stakes areas. In the present work, we establish how and when hallucinations arise in pre-trained transformer models through concept representations captured by sparse autoencoders, under scenarios with experimentally controlled uncertainty in the input space. Our systematic experiments reveal that the number of semantic concepts used by the transformer model grows as the input information becomes increasingly unstructured. In the face of growing uncertainty in the input space, the transformer model becomes prone to activate coherent yet input-insensitive semantic features, leading to hallucinated output. At its extreme, for pure-noise inputs, we identify a wide variety of robustly triggered and meaningful concepts in the intermediate activations of pre-trained transformer models, whose functional integrity we confirm through targeted steering. We also show that hallucinations in the output of a transformer model can be reliably predicted from the concept patterns embedded in transformer layer activations. This collection of insights on transformer internal processing mechanics has immediate consequences for aligning AI models with human values, AI safety, opening the attack surface for potential adversarial attacks, and providing a basis for automatic quantification of a model's hallucination risk.

幻觉机制模型解释AI安全自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。