深度越深,梯度下降越能逼近稀疏解,对初始值更不敏感。
Linear regression with overparameterized linear neural networks: Tight upper and lower bounds for implicit $\ell^1$-regularization
- 用深度大于等于2的线性网络研究过参数回归中的隐式正则化
- 深度≥3时误差随初始化尺度线性下降,深度=2时按α^{1−ϱ}下降
- 理论证明紧致边界,为深层网络泛化优势提供解释
现代机器学习模型常处于参数量超过样本数的过参数化状态。为理解梯度下降在此类模型中的隐式偏差,已有研究聚焦于深度D≥2的对角线线性神经网络在回归问题中的表现。当权重小初始化时,梯度下降倾向于选择ℓ¹-范数最小的解,即产生隐式ℓ¹正则化。本文分析了梯度流轨迹极限点与ℓ¹最小化解之间的近似误差,并推导出该误差的紧致上下界。结果表明:当D≥3时,误差随初始化尺度α线性减小;而D=2时,误差以α^{1−ϱ}速率下降,其中参数ϱ∈[0,1)可显式刻画,且与稀疏恢复中的零空间性质常数密切相关。通过具体例子验证了边界的渐近紧致性,数值实验支持理论结论,表明深层网络(D≥3)在合理初始化下可能具有更好泛化性能。
原文摘要 · Abstract (English)
Modern machine learning models are often trained in a setting where the number of parameters exceeds the number of training samples. To understand the implicit bias of gradient descent in such overparameterized models, prior work has studied diagonal linear neural networks in the regression setting. These studies have shown that, when initialized with small weights, gradient descent tends to favor solutions with minimal $\ell^1$-norm - an effect known as implicit regularization. In this paper, we investigate implicit regularization in diagonal linear neural networks of depth $D\ge 2$ for overparameterized linear regression problems. We focus on analyzing the approximation error between the limit point of gradient flow trajectories and the solution to the $\ell^1$-minimization problem. By deriving tight upper and lower bounds on the approximation error, we precisely characterize how the approximation error depends on the scale of initialization $α$. Our results reveal a qualitative difference between depths: for $D \ge 3$, the error decreases linearly with $α$, whereas for $D=2$, it decreases at rate $α^{1-\varrho}$, where the parameter $\varrho \in [0,1)$ can be explicitly characterized. Interestingly, this parameter is closely linked to so-called null space property constants studied in the sparse recovery literature. We demonstrate the asymptotic tightness of our bounds through explicit examples. Numerical experiments corroborate our theoretical findings and suggest that deeper networks, i.e., $D \ge 3$, may lead to better generalization, particularly for realistic initialization scales.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。