arXiv:2510.15804cs.CL2025-10NeurIPS被引 12

揭示大模型中真假语句线性分离的形成机制

Emergence of Linear Truth Encodings in Language Models

  • 用单层Transformer模型模拟真相编码的生成过程
  • 发现模型在事实共现数据中逐步学会线性区分真假
  • 适合研究语言模型内在表征与推理机制的读者

近期探测研究表明,大型语言模型存在分离真伪陈述的线性子空间,但其形成机制尚不明确。本文引入一个透明的一层Transformer玩具模型,能够端到端再现此类真相子空间,并揭示其产生的具体路径。研究一种简单场景:事实性陈述与其它事实性陈述共现(非事实亦然),促使模型为降低未来词的语言建模损失而学习该区分。实验验证了预训练模型中存在此模式。在玩具模型中观察到两阶段学习动态:初期快速记忆个别事实关联,随后在更长周期内学习线性分离真伪,从而降低语言建模损失。结果共同提供了机制性演示和实证依据,说明线性真相表征如何且为何在语言模型中出现。

原文摘要 · Abstract (English)

Recent probing studies reveal that large language models exhibit linear subspaces that separate true from false statements, yet the mechanism behind their emergence is unclear. We introduce a transparent, one-layer transformer toy model that reproduces such truth subspaces end-to-end and exposes one concrete route by which they can arise. We study one simple setting in which truth encoding can emerge: a data distribution where factual statements co-occur with other factual statements (and vice-versa), encouraging the model to learn this distinction in order to lower the LM loss on future tokens. We corroborate this pattern with experiments in pretrained language models. Finally, in the toy setting we observe a two-phase learning dynamic: networks first memorize individual factual associations in a few steps, then -- over a longer horizon -- learn to linearly separate true from false, which in turn lowers language-modeling loss. Together, these results provide both a mechanistic demonstration and an empirical motivation for how and why linear truth representations can emerge in language models.

语言模型真相编码线性子空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。