arXiv:2602.10496cs.LGcs.AI2026-02被引 7

Transformer训练轨迹快速压缩到3-4维子空间,揭示核心计算机制。

Low-Dimensional Execution Manifolds in Transformer Learning Dynamics: Evidence from Modular Arithmetic Tasks

  • 通过模运算任务发现参数轨迹坍缩至3-4维执行流形。
  • 92%的非交换性集中在正交方向,早期优化方向与流形对齐达10倍基准。
  • 适合关注模型可解释性、过参数化本质的研究者。

我们通过精心设计的模运算任务,研究了过参数化Transformer模型的学习动态几何结构。主要发现是:尽管参数空间维度高达128,训练轨迹仍迅速坍缩至3–4维的执行流形。该坍缩现象在不同随机种子和中等难度任务下均稳定存在,但流形在参数空间中的朝向随运行变化。这一几何结构解释了多个实证现象:(1) 注意力集中表现为流形内路由坐标的饱和;(2) SGD换位子早期与执行子空间对齐(达随机基线10倍),超过92%的非交换性被限制在正交方向,随训练收敛而减弱;(3) 稀疏自编码器仅捕捉辅助路由结构,无法分离出执行本身,后者仍分布于低维流形。结果表明,大部分参数用于抑制优化干扰,而核心计算发生在显著缩小的子空间中。这些发现为理解Transformer学习提供了统一的几何框架,对可解释性、训练课程设计及过参数化作用有重要意义。

原文摘要 · Abstract (English)

We investigate the geometric structure of learning dynamics in overparameterized transformer models through carefully controlled modular arithmetic tasks. Our primary finding is that despite operating in high-dimensional parameter spaces ($d=128$), transformer training trajectories rapidly collapse onto low-dimensional execution manifolds of dimension $3$--$4$. This dimensional collapse is robust across random seeds and moderate task difficulties, though the orientation of the manifold in parameter space varies between runs. We demonstrate that this geometric structure underlies several empirically observed phenomena: (1) sharp attention concentration emerges as saturation along routing coordinates within the execution manifold, (2) SGD commutators are preferentially aligned with the execution subspace (up to $10\times$ random baseline) early in training, with $>92\%$ of non-commutativity confined to orthogonal staging directions and this alignment decreasing as training converges, and (3) sparse autoencoders capture auxiliary routing structure but fail to isolate execution itself, which remains distributed across the low-dimensional manifold. Our results suggest a unifying geometric framework for understanding transformer learning, where the vast majority of parameters serve to absorb optimization interference while core computation occurs in a dramatically reduced subspace. These findings have implications for interpretability, training curriculum design, and understanding the role of overparameterization in neural network learning.

Transformer学习动态几何结构可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。