解释大模型如何通过预测下一个词学会人类可理解的概念。
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
- 用潜在离散变量建模概念,通过生成过程揭示表示机制
- 证明语言模型学得的表征近似于概念后验概率的对数
- 为稀疏自编码器评估提供理论依据,适合理论研究者
近期实证研究表明,大语言模型的表征包含人类可理解的概念,但这些表征的形成机制仍不明确。为此,我们提出一种新型生成模型,基于以潜在离散变量形式表达的概念生成文本。在弱条件下,即使潜在空间到观测空间的映射不可逆,我们建立了严格的可识别性结果:通过下一个词预测学习到的表征,可近似表示为给定输入上下文下这些潜在概念后验概率的对数,仅受线性变换影响。该理论发现表明:1)大语言模型捕捉到了底层生成因素;2)为线性表征假设提供了统一且严谨的视角;3)启发了一种有理论基础的稀疏自编码器评估方法。我们在模拟数据及Pythia、Llama和DeepSeek模型家族上验证了理论结果。
原文摘要 · Abstract (English)
Recent empirical evidence shows that LLM representations encode human-interpretable concepts. Nevertheless, the mechanisms by which these representations emerge remain largely unexplored. To shed further light on this, we introduce a novel generative model that generates tokens on the basis of such concepts formulated as latent discrete variables. Under mild conditions, even when the mapping from the latent space to the observed space is non-invertible, we establish rigorous identifiability result: the representations learned by LLMs through next-token prediction can be approximately modeled as the logarithm of the posterior probabilities of these latent discrete concepts given input context, up to an linear transformation. This theoretical finding: 1) provides evidence that LLMs capture essential underlying generative factors, 2) offers a unified and principled perspective for understanding the linear representation hypothesis, and 3) motivates a theoretically grounded approach for evaluating sparse autoencoders. Empirically, we validate our theoretical results through evaluations on both simulation data and the Pythia, Llama, and DeepSeek model families.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。