arXiv:2607.16741cs.LG2026-07

小模型中真理信号的维度随知识变化,可被分解为特定组件贡献。

The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models

论文配图:The Anatomy of a Truth Direction: Knowledge-Dependent Dimensionality, a Relational Law, and a Convergent Category Geometry in Small Language Models
图 1 · 摘自论文原文
  • 通过隐藏状态最小对的SVD探测真理方向,发现其维度依赖于模型知识水平
  • 注意力传播真理帧,前馈网络抵消当前块信号,SwiGLU值流导致峰值后衰减
  • 不同模型类别真理轴具语义符号排列,经谱共识修正后形成知识门控规律

Bürger等人(2024)表明,大语言模型中的真理表示在陈述极性上具有普适性,但存在于多维子空间中。我们沿三个问题展开:真理子空间维度如何随模型知识变化、哪个结构组件构建真理方向、该方向由何构成。第一部分,基于SVD的无训练方向探测显示,已知事实的真理信号集中于单一轴,知识减少时则扩散。第二部分,跨多个模型家族发现一种关系律:注意力传播真理帧,前馈网络抵消当前块帧,峰值后衰减由SwiGLU值流因果驱动;按类别形成的真理轴呈现语义符号排列,并在不同家族间收敛。压力测试揭示方向符号不稳定性,我们通过谱共识度量修复,使收敛成为知识门控律。最后,在Gemma-2-2b上的复现实验扩展分解工具以适配其夹心归一化,验证了这些规律与归因。我们量化了知识门为经典衰减,并分离出稳定、模型特异的私有几何结构。

原文摘要 · Abstract (English)

Bürger et al. (2024) demonstrated that truth representations in large language models are universal across statement polarity but reside within a multidimensional subspace. We extend this framework along three questions: how the dimensionality of the subspace depends on the model's knowledge, which architectural component builds the truth direction, and what the direction is a mixture of. In Part I, a training-free directional probe derived from the SVD of hidden-state minimal pairs shows that the dimensionality of truth is knowledge-dependent: the signal concentrates on a single axis for known facts and diffuses as knowledge decreases. In Part II, a relational law emerges across multiple model families: attention propagates truth frames, the feed-forward network opposes the current block's frame, and post-peak decay is causally attributed to the SwiGLU value stream. Furthermore, per-category truth axes form a semantically signed arrangement that converges across families. Stress tests expose a sign instability in this orientation, which we repair with a spectral consensus gauge to sharpen the convergence into a knowledge-gated law. Finally, a replication campaign on Gemma-2-2b, extending our decomposition tools to accommodate its sandwich normalization, confirms these laws and attributions. We quantify the knowledge gate as classical attenuation and isolate a stable, model-specific private geometry.

语言模型真理方向知识依赖几何结构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。