在自然文本中植入可调控隐变量,验证大模型如何跟踪信念状态及其概念几何。
Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry

- 在自然文本中嵌入8个无关的稀疏自编码器方向作为隐变量
- 小模型成功追踪到隐变量的后验分布,并将其排列成环形结构
- 首次将信念状态与概念几何的统计动力学联系起来,适合大模型可解释性研究者
大语言模型被认为会追踪‘信念状态’,即对控制语言的隐变量进行概率推断(Shai等,2024;Sarfati等,2026),但此前仅在合成数据和零星案例中被证实。且从未与模型特征的几何结构(概念可解释性所发现的激活模式)建立实证关联。本文在自然文本中植入可控的隐变量:由一个大语言模型教师生成普通文本时,我们以子阈值方式在每个词元上沿K=8个不相关的稀疏自编码器方向之一进行引导,活跃方向遵循环形马尔可夫链。训练一个小的Transformer模型后,该模型确能追踪所植入隐变量的贝叶斯后验信念。此外,它还将8个状态以精确的马尔可夫链顺序排列成环状,为概念几何可能由其背后隐变量的统计动态形成提供了支持证据。
原文摘要 · Abstract (English)
LLMs are thought to track "belief states," i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sarfati et al., 2026), but so far this has only been comprehensively demonstrated on toy synthetic data and in a few isolated case studies. It has also never been empirically connected to the geometry of LLM features (the concepts interpretability finds in model activations). In this work, we plant a controllable latent variable inside natural-looking text. An LLM teacher writes ordinary text while we "subliminally" steer it along one of K = 8 unrelated sparse autoencoder directions at each token, with the active directions following a ring-shaped Markov chain. A small transformer model trained on this corpus does indeed track the Bayesian posterior belief about our planted latent variable. Moreover, it also arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。