研究深度网络梯度下降的隐式偏置,发现训练后模型会自动最大化几何间隔。
Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural Networks
- 用极小学习率的最速下降法优化深层网络
- 训练达100%准确后几何间隔持续增长
- 结果解释了Adam等自适应方法的泛化优势
我们研究了在深度齐次神经网络中,以无穷小学习率运行的广义最速下降算法的隐式偏置。结果显示:(a) 一旦网络达到完美训练精度,算法相关的几何间隔便开始增加;(b) 训练轨迹的任意极限点对应于相应最大间隔优化问题的KKT点。我们通过实验分析了多种最速下降算法下网络的优化路径,揭示了其与流行自适应方法(如Adam和Shampoo)隐式偏置之间的关联。
原文摘要 · Abstract (English)
We study the implicit bias of the general family of steepest descent algorithms with infinitesimal learning rate in deep homogeneous neural networks. We show that: (a) an algorithm-dependent geometric margin starts increasing once the networks reach perfect training accuracy, and (b) any limit point of the training trajectory corresponds to a KKT point of the corresponding margin-maximization problem. We experimentally zoom into the trajectories of neural networks optimized with various steepest descent algorithms, highlighting connections to the implicit bias of popular adaptive methods (Adam and Shampoo).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。