学习率调控模型范数与尖锐度的权衡,影响泛化性能。
Conflicting Biases at the Edge of Stability: Norm versus Sharpness Regularization
- 通过实验发现学习率在控制模型范数和尖锐度间起关键调节作用。
- 单一依赖范数或尖锐度正则化都无法最小化泛化误差。
- 揭示了训练中范数与尖锐度动态博弈的机制,适合研究优化隐性偏置者。
过参数化网络的出色泛化能力常归因于隐性偏差,如小学习率下的范数最小化,以及边缘稳定性状态下的低尖锐度。本文认为,理解梯度下降的泛化性能需分析这些不同形式隐性正则化的相互作用。我们通过实验证明,学习率在训练模型的低参数范数与低尖锐度之间起插值作用。此外,我们证明在简单回归任务中,对角线线性网络仅依靠任一隐性偏差均无法最小化泛化误差。这些发现表明,仅关注单一隐性偏差不足以解释良好泛化,从而推动更全面的隐性正则化视角,以捕捉非可忽略学习率下范数与尖锐度之间的动态权衡。
原文摘要 · Abstract (English)
The remarkable generalization properties of overparameterized networks are often attributed to implicit biases, such as norm minimization at small learning rates and low sharpness in the Edge-of-Stability regime. In this work, we argue that a comprehensive understanding of the generalization performance of gradient descent requires analyzing the interaction between these various forms of implicit regularization. We empirically demonstrate that the learning rate interpolates between low parameter norm and low sharpness of the trained model. We furthermore prove that neither implicit bias alone minimizes the generalization error for diagonal linear networks trained on a simple regression task. These findings demonstrate that focusing on a single implicit bias is insufficient to explain good generalization, and they motivate a broader view of implicit regularization that captures the dynamic trade-off between norm and sharpness induced by non-negligible learning rates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。