arXiv:2604.04655cs.LGcond-mat.dis-nn2026-04被引 4

神经网络从记诵到泛化的突变,本质是维度相变。

Grokking as Dimensional Phase Transition in Neural Networks

  • 用梯度雪崩动力学分析模型规模,发现泛化始于有效维度跃迁。
  • 维度从低于1跃升至高于1,对应学习从过拟合到泛化的临界点。
  • 该现象受反向传播相关性影响,与网络结构无关,适合研究大模型训练

神经网络的突现泛化现象——即从记忆数据到泛化能力的突然转变——挑战了我们对学习动态的理解。通过对八个不同规模模型的梯度雪崩动力学进行有限尺寸标度分析,我们发现,这种突现泛化本质上是一种 extit{维度相变}:有效维度~$D$在泛化起始时刻从亚扩散(亚临界,$D < 1$)跃迁至超扩散(超临界,$D > 1$),表现出自组织临界性(SOC)。关键的是,$D$反映的是梯度场几何结构,而非网络架构:合成的独立同分布高斯梯度即使在不同拓扑下也维持$D ightarrow 1$;而真实训练中因反向传播相关性导致维度过剩。这一在多种拓扑下均稳健出现的$D(t)$跨阈现象,为理解过参数化网络的可训练性提供了新视角。

原文摘要 · Abstract (English)

Neural network grokking -- the abrupt memorization-to-generalization transition -- challenges our understanding of learning dynamics. Through finite-size scaling of gradient avalanche dynamics across eight model scales, we find that grokking is a \textit{dimensional phase transition}: effective dimensionality~$D$ crosses from sub-diffusive (subcritical, $D < 1$) to super-diffusive (supercritical, $D > 1$) at generalization onset, exhibiting self-organized criticality (SOC). Crucially, $D$ reflects \textbf{gradient field geometry}, not network architecture: synthetic i.i.d.\ Gaussian gradients maintain $D \approx 1$ regardless of graph topology, while real training exhibits dimensional excess from backpropagation correlations. The grokking-localized $D(t)$ crossing -- robust across topologies -- offers new insight into the trainability of overparameterized networks.

神经网络泛化相变训练动态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。