揭示深度ReLU网络最小范数插值的稳定条件。
Sufficient Conditions for Stability of Minimum-Norm Interpolating Deep ReLU Networks
- 通过构建稳定子网络+低秩层的结构,保证算法稳定性。
- 若后续层非低秩,则即使有稳定子网络也不稳定。
- 适合研究模型泛化与深度学习优化机制的研究者。
算法稳定性是分析学习算法泛化误差的经典框架,认为当算法对训练集的小扰动(如删除或替换一个样本)不敏感时,其泛化误差较小。尽管该框架已成功应用于多种经典算法,但在深度神经网络分析中进展有限。本文研究了实现零训练误差的深度ReLU同质网络在最小L2范数插值下的算法稳定性。发现:1)当网络包含一个(可能很小的)稳定子网络,且其后接一层低秩权重矩阵时,网络具有稳定性;2)即使存在稳定子网络,若后续层非低秩,则网络不一定稳定。低秩假设受到近期实证和理论研究支持,表明在最小范数插值和权重衰减正则化下,深度神经网络训练倾向于产生低秩权重矩阵。
原文摘要 · Abstract (English)
Algorithmic stability is a classical framework for analyzing the generalization error of learning algorithms. It predicts that an algorithm has small generalization error if it is insensitive to small perturbations in the training set such as the removal or replacement of a training point. While stability has been demonstrated for numerous well-known algorithms, this framework has had limited success in analyses of deep neural networks. In this paper we study the algorithmic stability of deep ReLU homogeneous neural networks that achieve zero training error using parameters with the smallest $L_2$ norm, also known as the minimum-norm interpolation, a phenomenon that can be observed in overparameterized models trained by gradient-based algorithms. We investigate sufficient conditions for such networks to be stable. We find that 1) such networks are stable when they contain a (possibly small) stable sub-network, followed by a layer with a low-rank weight matrix, and 2) such networks are not guaranteed to be stable even when they contain a stable sub-network, if the following layer is not low-rank. The low-rank assumption is inspired by recent empirical and theoretical results which demonstrate that training deep neural networks is biased towards low-rank weight matrices, for minimum-norm interpolation and weight-decay regularization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。