arXiv:2505.21399cs.CLcs.AI2025-05被引 2

大模型生成时能自我判断事实正确性,内部有稳定信号支持。

Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling

  • 通过残差流发现模型内存在决定事实正确性的线性特征。
  • 该自知信号对格式微变鲁棒,且在中层结构达到峰值。
  • 适合关注模型可解释性与可靠性的人阅读。

大语言模型生成内容中的事实错误是其广泛应用的主要担忧之一。以往研究发现,语言模型有时能在生成后检测到事实错误(即事后事实核查)。本文提供证据表明,语言模型在生成时刻已具备内在的正确性判断能力。我们证明,对于特定主体实体与关系,模型在Transformer残差流中编码了线性特征,决定其能否正确回忆属性(构成有效实体-关系-属性三元组)。该自知信号对轻微格式变化具有鲁棒性。通过不同示例选择策略考察上下文扰动的影响,结果表明:在模型规模与训练动态的缩放实验中,自知能力在训练初期迅速出现,并在中间层达到峰值。这些发现揭示了大语言模型中固有的自我监控能力,有助于提升其可解释性与可靠性。

原文摘要 · Abstract (English)

Factual incorrectness in generated content is one of the primary concerns in ubiquitous deployment of large language models (LLMs). Prior findings suggest LLMs can (sometimes) detect factual incorrectness in their generated content (i.e., fact-checking post-generation). In this work, we provide evidence supporting the presence of LLMs' internal compass that dictate the correctness of factual recall at the time of generation. We demonstrate that for a given subject entity and a relation, LLMs internally encode linear features in the Transformer's residual stream that dictate whether it will be able to recall the correct attribute (that forms a valid entity-relation-attribute triplet). This self-awareness signal is robust to minor formatting variations. We investigate the effects of context perturbation via different example selection strategies. Scaling experiments across model sizes and training dynamics highlight that self-awareness emerges rapidly during training and peaks in intermediate layers. These findings uncover intrinsic self-monitoring capabilities within LLMs, contributing to their interpretability and reliability.

自知能力可解释性模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。