arXiv:2605.14258cs.LGcs.AI2026-05被引 6

揭示大模型残差流中谱几何与网络拓扑的动态耦合机制

Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology

论文配图:Dynamics of the Transformer Residual Stream: Coupling Spectral Geometry to Network Topology
图 1 · 摘自论文原文
  • 通过全雅可比特征分解,发现训练使谱梯度沿深度单调变化
  • 早期层具旋转主导的非对称性,晚期层趋近对称,存在低秩瓶颈
  • 该结构是学习所得,与功能拓扑相关,适合研究模型内部动力学

大型语言模型表现卓越,但其计算如何在层间传播仍不清晰。现有研究将深度视为离散时间,残差流视为动力系统,每层的非线性更新有局部线性描述。然而先前分析依赖标量摘要或近似线性化,未揭示训练后大模型的完整谱几何。我们对三个生产级大模型进行了全雅可比特征分解,发现训练引入了沿深度的单调谱梯度:从早期层的非正则、旋转主导型过渡到晚期层的近对称型,并伴随累积的低秩瓶颈,将扰动压缩至残差流有效维度的一小部分。实验表明,此梯度与维度坍缩是学习所得而非架构决定,当结构化非正则性被移除时基本消失。进一步发现,图社区的拓扑位置可预测雅可比是否放大或抑制其,耦合符号由局部算子类型决定,该关系在初始化时不存在。这些结果揭示了大模型中学习到的谱几何,将扰动传播与压缩关联到网络的功能拓扑。

原文摘要 · Abstract (English)

Large language models are remarkably capable, yet how computation propagates through their layers remains poorly understood. A growing line of work treats depth as discrete time and the residual stream as a dynamical system, where each layer's nonlinear update has a local linear description. However, previous analyses have relied on scalar summaries or approximate linearizations, leaving the full spectral geometry of trained LLMs unknown. We perform full Jacobian eigendecomposition across three production--scale LLMs and show that training installs a monotonic spectral gradient through depth -- from non-normal, rotation-dominated early layers to near--symmetric late layers -- together with a cumulative low-rank bottleneck that funnels perturbations into a small fraction of the residual stream's effective dimensions. Our experiments reveal that this gradient and the dimensional collapse are learned rather than architectural, and is largely dissolved when structured non-normality is removed. We further show that the topological positioning of graph communities predicts whether the Jacobian amplifies or suppresses them, with the sign of the coupling determined by the local operator type, a relationship absent at initialization. These results map a learned spectral geometry in LLMs that links perturbation propagation and compression to the network's functional topology.

Transformer残差流谱几何大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。