用双时间尺度理论解释SGD中的异常现象
A dynamic view of some anomalous phenomena in SGD
- 基于双时间尺度随机逼近理论分析梯度动态
- 揭示了测试误差先降后升再降的双下降现象
- 适合研究优化动态与泛化性能关系的读者
Belkin等人观察到,过参数化的神经网络会出现'双下降'现象:随着模型复杂度(以特征数衡量)增加,测试误差先下降,再上升,随后再次下降。在训练周期层面也存在类似现象——测试误差随迭代次数先降后升再降。另一个异常现象是'grokking',即两个下降阶段被一个均值损失几乎不变的第三阶段打断。本文通过将梯度动力学的连续时间极限与双时间尺度随机逼近理论结合,为这些现象提供了合理的解释,为已有研究主题提供了新视角。
原文摘要 · Abstract (English)
It has been observed by Belkin et al.\ that over-parametrized neural networks exhibit a `double descent' phenomenon. That is, as the model complexity (as reflected in the number of features) increases, the test error initially decreases, then increases, and then decreases again. A counterpart of this phenomenon in the time domain has been noted in the context of epoch-wise training, viz., the test error decreases with the number of iterates, then increases, then decreases again. Another anomalous phenomenon is that of \textit{grokking} wherein two regimes of descent are interrupted by a third regime wherein the mean loss remains almost constant. This note presents a plausible explanation for these and related phenomena by using the theory of two time scale stochastic approximation, applied to the continuous time limit of the gradient dynamics. This gives a novel perspective for an already well studied theme.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。