arXiv:2601.06599cs.CLcs.AI2026-01ACL

揭示上下文如何通过几何变化影响大模型对真假的表示

How Context Shapes Truth: Geometric Transformations of Statement-level Truth Representations in LLMs

  • 分析上下文引入后真值向量的方向与幅度变化
  • 上下文使真假表示在激活空间中更易区分,幅度普遍增大
  • 大模型依赖方向变化辨析相关上下文,小模型依赖幅度差异

大型语言模型(LLMs)常将语句真假编码为残差流激活中的向量,即真值向量。已有研究关注真值向量本身,但其在引入上下文后的动态变化尚不明确。本文通过测量(1)有无上下文时真值向量的方向变化(θ)与(2)加入上下文后真值向量的相对幅度,系统考察该变化。在四个主流LLM和四个数据集上发现:(1)早期层中真值向量近似正交,中层趋于收敛,后期可能稳定或继续增长;(2)引入上下文通常提升真值向量幅度,放大真假表示间的空间分离;(3)大模型主要通过方向变化区分相关与无关上下文,小模型则依赖幅度差异。此外,与参数化知识冲突的上下文比一致上下文引发更大的几何变化。这些结果共同提供了上下文如何重塑模型内部真值表示的几何刻画。

原文摘要 · Abstract (English)

Large Language Models (LLMs) often encode whether a statement is true as a vector in their residual stream activations. These vectors, also known as truth vectors, have been studied in prior work, however how they change when context is introduced remains unexplored. We study this question by measuring (1) the directional change ($θ$) between the truth vectors with and without context and (2) the relative magnitude of the truth vectors upon adding context. Across four LLMs and four datasets, we find that (1) truth vectors are roughly orthogonal in early layers, converge in middle layers, and may stabilize or continue increasing in later layers; (2) adding context generally increases the truth vector magnitude, i.e., the separation between true and false representations in the activation space is amplified; (3) larger models distinguish relevant from irrelevant context mainly through directional change ($θ$), while smaller models show this distinction through magnitude differences. We also find that context conflicting with parametric knowledge produces larger geometric changes than parametrically aligned context. Collectively, these findings provide a geometric characterization of how context transforms the truth vector in the activation space of LLMs.

大模型机制真值表示上下文影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。