发现神经网络从记忆到泛化的突变中存在维度临界现象。
Dimensional Criticality at Grokking Across MLPs and Transformers
- 用梯度快照构建级联模型,提取时间分辨的级联维数D(t)
- 在泛化突变点处,D(t)精确穿过高斯扩散基准值1,方向依任务而异
- 该现象具可预测性,提前100-200轮即与未泛化轨迹分离
复杂系统中不同动力学态间的突变是典型特征。深度神经网络中的grocking现象尤为突出——在训练准确率饱和后,突然从记忆转向泛化。然而这种转变的宏观表征仍不清晰。本文提出TDU-OFC(阈值扩散更新-奥拉米-费德-克里斯滕森)离线级联探测器,将梯度快照转化为级联统计量,并通过与grocking对齐的有限尺寸标度,提取出一个宏观可观测量:时间分辨的有效级联维数D(t)。在模块加法任务的Transformer和异或任务的MLP上,我们发现该维数在泛化转变点恰好穿过高斯扩散基准值D=1,且方向由任务决定:模块加法下降穿越(从D>1),异或上升穿越(从D<1)。这种反向收敛行为支持存在一个共享的临界流形,而非偶然停留在D≈1附近。负控制实验表明,未发生grocking的训练始终处于超临界态(D>1),从未进入后转变区间。此外,级联分布呈现重尾特性,且有限尺寸标度与从D(t)提取的维度指数一致。影子探测控制(α_train=0)证实D(t)非侵入性,且已泛化轨迹在行为转变前100–200个周期即开始偏离未泛化轨迹。
原文摘要 · Abstract (English)
Abrupt transitions between distinct dynamical regimes are a hallmark of complex systems. Grokking in deep neural networks provides a striking example -- an abrupt transition from memorization to generalization long after training accuracy saturates -- yet robust macroscopic signatures of this transition remain elusive. Here we introduce \textbf{TDU--OFC} (Thresholded Diffusion Update--Olami-Feder-Christensen), an offline avalanche probe that converts gradient snapshots into cascade statistics and extracts a \emph{macroscopic observable} -- the time-resolved effective cascade dimension $D(t)$ -- via grokking-aligned finite-size scaling. Across Transformers trained on modular addition and MLPs trained on XOR, we discover a localized dynamical crossing of the Gaussian diffusion baseline $D=1$ precisely at the generalization transition. The crossing direction is task-dependent: modular addition descends through $D=1$ (approaching from $D>1$), while XOR ascends (from $D<1$). This opposite-direction convergence is consistent with attraction toward a candidate shared critical manifold, rather than trivial residence near $D \approx 1$. Negative controls confirm this picture: ungrokked runs remain supercritical ($D>1$) and never enter the post-transition regime. In addition, avalanche distributions exhibit heavy tails and finite-size scaling consistent with the dimensional exponent extracted from $D(t)$. Shadow-probe controls ($α_{\mathrm{train}}=0$) confirm that $D(t)$ is non-invasive, and grokked trajectories diverge from ungrokked ones in $D(t)$ some $100$--$200$ epochs before the behavioral transition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。