arXiv:2608.25813cs.LG2026-08

通过脉冲权重衰减,发现模型在泛化前已出现功能选择的稳定路径。

Canalization Before Generalization: Grokking as a Dynamical Probe

论文配图:Canalization Before Generalization: Grokking as a Dynamical Probe
图 1 · 摘自论文原文
  • 用短时权重衰减脉冲探测训练中期的泛化潜力变化。
  • 泛化时间随脉冲强度呈现稳定先后顺序,早于可见泛化出现。
  • 揭示了模型在训练中逐步锁定最优解的动态机制,适合研究泛化形成过程。

对于过参数化的神经网络,存在多个能完美拟合训练数据但对未见样本表现迥异的解。Grokking 现象将训练拟合与可见泛化分离开来,为研究这一选择如何在训练过程中发展提供了窗口。我们在三个 grokking 任务中,在训练平台期施加短时、固定时长的权重衰减(WD)脉冲,并测量其对后续泛化时间的影响。早期脉冲导致的泛化时间变化无序,后期则形成稳定的剂量排序:更强的权重衰减使泛化更早发生,更强的抑制则延后泛化。该排序在所有三个任务中均先于可见泛化出现。与此同时,扰动解与基准解之间的测试损失壁垒逐渐趋近于零,而剂量有序的时间效应仍持续存在。我们称这种解决方案选择日益受限且剂量依赖的时间敏感性持续存在的现象为函数选择的渠道化(canalization)。

原文摘要 · Abstract (English)

For overparameterized neural networks, many solutions can fit the training data equally well while behaving very differently on unseen samples. Grokking separates training fit from visible generalization, providing a window for studying how this selection develops during training. We sweep short, fixed-duration weight-decay (WD) pulses across this plateau and measure how they shift later generalization time. Across three grokking tasks, these shifts are unordered early in the plateau but later form a stable dose ordering, with stronger WD increases leading to earlier generalization and stronger WD decreases leading to later generalization. This ordering emerges before visible generalization in all three tasks. Meanwhile, test-loss barriers between perturbed and baseline generalization checkpoints collapse toward zero while the ordered timing effects persist. We call this combination of increasingly constrained solution selection and persistent dose-ordered timing sensitivity the canalization of function selection.

泛化机制模型动力学权重衰减函数选择

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。