arXiv:2601.19942cs.LGcs.CL2026-01被引 2

发现大模型推理能力突现的几何机制,揭示隐藏状态的相变规律。

Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds

  • 从统计物理视角分析深层Transformer的隐状态流形,用谱分析捕捉演化特征。
  • 在约0.42归一化深度处观察到有效维度骤降,对应推理能力跃迁的关键点。
  • 提出可复用的瞬态概念结构(TCO),适合研究模型内部表征与逻辑推理的人参考。

我们从几何与统计物理角度研究深层Transformer语言模型中多步推理能力的涌现。将隐状态轨迹视为隐式黎曼流形上的流动,分析各层激活的协方差谱 $C^{(\ ext{ℓ})}=\mathbb{E}[h^{(\text{ℓ})}h^{(\text{ℓ})\top}]$,并追踪其与随机矩阵本体的偏差。在1.5B至30B参数规模模型中,观察到有效维度显著降低,符合相变特征:基于稀疏性/局域化的序参量 $Ω(h)=1-\|h\|_1/(\sqrt{d}\|h\|_2)$ 在足够大模型中于临界归一化深度 $γ_c\approx 0.42$ 处出现不连续。我们将前向传播形式化为离散粗粒化映射,将稳定‘概念洼地’与该重整化类动力学的固定点关联。结果低熵态表现为谱尾塌陷及表示空间中形成瞬时、可复用的对象结构,称为瞬态类别对象(TCOs)。通过多组开源模型的层间探测验证了理论预测的特征。

原文摘要 · Abstract (English)

We study the emergence of multi-step reasoning in deep Transformer language models through a geometric and statistical-physics lens. Treating the hidden-state trajectory as a flow on an implicit Riemannian manifold, we analyze the layerwise covariance spectrum of activations, where $C^{(\ell)}=\mathbb{E}[h^{(\ell)}h^{(\ell)\top}]$, and track deviations from a random-matrix bulk. Across model scales (1.5B--30B), we observe a sharp reduction in effective dimensionality consistent with a phase transition: an order parameter based on sparsity/localization, $Ω(h)=1-\|h\|_1/(\sqrt{d}\|h\|_2)$, exhibits a discontinuity near a critical normalized depth $γ_c\approx 0.42$ in sufficiently large models. We formalize the forward pass as a discrete coarse-graining map and relate the appearance of stable "concept basins" to fixed points of this renormalization-like dynamics. The resulting low-entropy regime is characterized by a spectral tail collapse and by the formation of transient, reusable object-like structures in representation space, which we call Transient Class Objects (TCOs). We provide theoretical conditions connecting logical separability to spectral decay and validate the predicted signatures with layerwise probes on multiple open-weight model families.

Transformer表征学习相变几何建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。