发现特征学习强度存在最优值,过强或过弱都影响模型泛化。
Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in Generalization
- 通过初始化尺度控制特征学习强度,研究其对训练的影响。
- 实验证明存在最优特征学习强度,能显著提升泛化性能。
- 揭示过强导致过度对齐、过弱引发过拟合的权衡机制。
特征学习强度(FLS),即模型有效输出缩放的倒数,在神经网络优化动态中起关键作用。尽管已有研究广泛探讨了在渐近情形下(训练时间与FLS)FLS的影响,但现有理论对实际训练场景中FLS如何影响泛化仍缺乏深入理解,例如当训练在达到目标训练风险时停止时。本文研究了在实际条件下深度网络中FLS对泛化的影响。通过实验发现,存在一个‘最优FLS’——既不太小也不太大——可带来显著的泛化增益,这与普遍认为更强特征学习总是改善泛化的直觉相悖。为解释该现象,我们对使用逻辑损失训练的两层ReLU网络中的梯度流动态进行了理论分析,其中通过初始化尺度控制FLS。主要理论结果表明,最优FLS的存在源于两种竞争效应的权衡:过大的FLS会引发‘过度对齐’现象,损害泛化;而过小的FLS则导致‘过拟合’。
原文摘要 · Abstract (English)
Feature learning strength (FLS), i.e., the inverse of the effective output scaling of a model, plays a critical role in shaping the optimization dynamics of neural nets. While its impact has been extensively studied under the asymptotic regimes -- both in training time and FLS -- existing theory offers limited insight into how FLS affects generalization in practical settings, such as when training is stopped upon reaching a target training risk. In this work, we investigate the impact of FLS on generalization in deep networks under such practical conditions. Through empirical studies, we first uncover the emergence of an $\textit{optimal FLS}$ -- neither too small nor too large -- that yields substantial generalization gains. This finding runs counter to the prevailing intuition that stronger feature learning universally improves generalization. To explain this phenomenon, we develop a theoretical analysis of gradient flow dynamics in two-layer ReLU nets trained with logistic loss, where FLS is controlled via initialization scale. Our main theoretical result establishes the existence of an optimal FLS arising from a trade-off between two competing effects: An excessively large FLS induces an $\textit{over-alignment}$ phenomenon that degrades generalization, while an overly small FLS leads to $\textit{over-fitting}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。