揭示大模型信念在表示空间中的几何结构与动态演化规律。
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
- 用曲面流形建模语言模型的隐含信念分布
- 线性干预会破坏信念结构导致意外耦合变化
- 几何感知干预可保持信念家族完整性,适合可控生成
大语言模型(LLMs)从提示中隐式形成关于潜在变量的信念(后验分布),但缺乏对这些信念如何在表示空间中编码、如何随新证据更新以及干预如何重塑它们的机制理解。我们研究了在上下文样本中由 Llama-3.2 推断正态分布参数的受控场景。结果表明,参数后验以表示空间中的弯曲流形形式存在,并可追踪其在提示中的演化过程。标准线性引导会将表示移出流形,引发未预期的耦合变化;而几何感知方法能保留目标信念族。本工作展示了线性场探测(LFP)作为一种系统性方法,可用于铺砌数据流形并实现尊重底层几何结构的干预。结果表明,LLM 的信念本质上是几何对象,全局线性表示常为不当抽象。
原文摘要 · Abstract (English)
Large language models (LLMs) form implicit beliefs (posteriors over latent variables) from prompts, but we lack a mechanistic account of how these beliefs are encoded in representation space, how they update with new evidence, and how interventions reshape them. We study a controlled setting in which Llama-3.2 infers the parameters of a normal distribution from in-context samples. We show that parameter posteriors are encoded as curved manifolds in representation space, and trace how they evolve along the prompt. Standard linear steering moves representations off-manifold, inducing unintended, coupled changes, whereas geometry-aware methods preserve the target belief family. Our work demonstrates an example of linear field probing (LFP) as a principled approach to tile the data manifold and make interventions that respect the underlying geometry. Our results suggest that LLM beliefs are inherently geometric objects, and that globally linear representations are often inadequate abstractions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。