arXiv:2604.07380cs.LG2026-04被引 1

揭示模型在泛化前的梯度与权重衰减协同演化规律。

The Lifecycle of the Spectral Edge: From Gradient Learning to Weight-Decay Compression

  • 分离参数更新的谱边缘为梯度与权重衰减成分,分析其动态变化。
  • 发现泛化前存在两阶段生命周期:梯度主导到压缩轴形成,后者扰动平坦但移除影响巨大(>4000倍)。
  • 提出三类通用行为模式,由间隙流方程预测,适用于理解模型内在机制。

我们在两个序列任务(Dyck-1 和 SCAN)中分解了谱边缘——参数更新的格拉姆矩阵主方向——的梯度与权重衰减成分,观察其在泛化过程中的演变。发现一个清晰的双阶段生命周期:泛化前,谱边缘由梯度驱动且功能活跃;泛化时,梯度与权重衰减对齐,谱边缘成为压缩轴,表现为扰动平坦但消融关键(比随机方向影响大4000倍以上)。三种普适性类别(功能型、混合型、压缩型)出现,由间隙流方程预测。非线性探测显示信息被重新编码而非丢失(MLP $R^2=0.99$,线性 $R^2=0.86$),且泛化后移除权重衰减会逆转压缩但仍保持算法有效。

原文摘要 · Abstract (English)

We decompose the spectral edge -- the dominant direction of the Gram matrix of parameter updates -- into its gradient and weight-decay components during grokking in two sequence tasks (Dyck-1 and SCAN). We find a sharp two-phase lifecycle: before grokking the edge is gradient-driven and functionally active; at grokking, gradient and weight decay align, and the edge becomes a compression axis that is perturbation-flat yet ablation-critical (>4000x more impactful than random directions). Three universality classes emerge (functional, mixed, compression), predicted by the gap flow equation. Nonlinear probes show information is re-encoded, not lost (MLP $R^2=0.99$ where linear $R^2=0.86$), and removing weight decay post-grok reverses compression while preserving the algorithm.

模型机制泛化压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。