arXiv:2602.16746cs.LGcs.AI2026-02被引 10

发现模型在低维子空间中训练,横向上曲率积累预示泛化跃迁。

Low-Dimensional and Transversely Curved Optimization Dynamics in Grokking

  • 用主成分分析发现训练轨迹集中于低维执行子空间,单主成分占方差68%-83%。
  • 横向往外的曲率急剧增长,且先于泛化出现,时间差符合幂律关系。
  • 只有沿学习子空间运动才能实现泛化,增加曲率无法替代这一过程。

小规模算法任务中的记忆到泛化的延迟转变——即‘领悟’现象——仍缺乏理解。本文对变压器模型在模运算任务上的优化动态进行几何分析。注意力权重轨迹的主成分分析显示,训练主要在低维执行子空间内进行,单一主成分解释了68%至83%的轨迹方差。为探测损失曲面几何结构,我们测量梯度步长的交换子缺陷(非对易性),并将其投影到该学习子空间。结果表明,垂直于执行子空间的方向上曲率迅速增长,而轨迹始终基本被限制在该子空间内。重要的是,曲率增长始终早于泛化发生,其领先时间在不同学习率与超参数下均服从沟克时间尺度的幂律。因果干预实验表明,沿学习子空间的运动是领悟所必需的,而人为增加曲率缺陷则无效。上述结果支持一种几何解释:领悟反映的是从低维束缚与横向曲率累积的亚稳态中逃逸的过程。所有发现均在多个学习率、一个不同慢速训练阶段(lr=5e-5, wd=0.1, 3层)及三个随机种子下重复验证,尽管各阶段对齐动力学存在定量差异。因果实验进一步确认:正交梯度流是必要但不充分条件;抑制它会以单调剂量响应方式阻止泛化,而增强曲率缺陷无影响。

原文摘要 · Abstract (English)

Grokking -- the delayed transition from memorization to generalization in small algorithmic tasks -- remains poorly understood. We present a geometric analysis of optimization dynamics in transformers trained on modular arithmetic. PCA of attention weight trajectories reveals that training evolves predominantly within a low-dimensional execution subspace, with a single principal component capturing 68-83% of trajectory variance. To probe loss-landscape geometry, we measure commutator defects -- the non-commutativity of successive gradient steps -- and project them onto this learned subspace. We find that curvature grows sharply in directions orthogonal to the execution subspace while the trajectory remains largely confined to it. Importantly, curvature growth consistently precedes generalization across learning rates and hyperparameter regimes, with the lead time obeying a power law in the grokking timescale. Causal intervention experiments show that motion along the learned subspace is necessary for grokking, while artificially increasing curvature is insufficient. Together, these results support a geometric account in which grokking reflects escape from a metastable regime characterized by low-dimensional confinement and transverse curvature accumulation. All findings replicate across this learning-rate range, a qualitatively different slow regime (lr=5e-5, wd=0.1, 3 layers), and three random seeds, though alignment dynamics differ quantitatively between regimes. Causal intervention experiments establish that orthogonal gradient flow is necessary but not sufficient for grokking: suppressing it prevents generalization with a monotonic dose-response across four operations, while artificially boosting curvature defects has no effect.

Transformer优化几何领悟现象

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。