arXiv:2505.12182cs.CL2025-05被引 1

发现语言模型中编码真实性的神经元,可提升模型可信度。

Truth Neurons

  • 通过分析神经元激活,定位语言模型中的真实性编码单元。
  • 抑制这些神经元会降低模型在多个数据集上的真实性表现。
  • 该机制跨模型、跨数据集通用,适合安全可信研究者参考。

尽管语言模型在多种任务中表现出色,但其生成内容有时不真实。我们对模型内部真实性机制的理解有限,影响其可靠性与安全性。本文提出一种在神经元层面识别真实性表征的方法,发现语言模型存在不依赖具体主题的‘真实性神经元’。在不同规模的模型上实验均验证了该现象,且神经元分布模式与已有真实性几何研究一致。在TruthfulQA数据集上识别出的真实性神经元,其激活被抑制后,不仅在TruthfulQA上性能下降,其他基准测试也受影响,说明真实性机制并非特定于某数据集。研究为理解语言模型的真实性机制提供了新视角,并指明提升模型可信度的新方向。

原文摘要 · Abstract (English)

Despite their remarkable success and deployment across diverse workflows, language models sometimes produce untruthful responses. Our limited understanding of how truthfulness is mechanistically encoded within these models jeopardizes their reliability and safety. In this paper, we propose a method for identifying representations of truthfulness at the neuron level. We show that language models contain truth neurons, which encode truthfulness in a subject-agnostic manner. Experiments conducted across models of varying scales validate the existence of truth neurons, confirming that the encoding of truthfulness at the neuron level is a property shared by many language models. The distribution patterns of truth neurons over layers align with prior findings on the geometry of truthfulness. Selectively suppressing the activations of truth neurons found through the TruthfulQA dataset degrades performance both on TruthfulQA and on other benchmarks, showing that the truthfulness mechanisms are not tied to a specific dataset. Our results offer novel insights into the mechanisms underlying truthfulness in language models and highlight potential directions toward improving their trustworthiness and reliability.

真实性神经元分析大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。