提出首个完整非渐近理论,解释深度网络泛化与双下降现象。
Path Regularization: A Near-Complete and Optimal Nonasymptotic Generalization Theory for Multilayer Neural Networks and Double Descent Phenomenon
- 基于路径正则化构建通用泛化上界,无需假设损失有界
- 首次揭示双下降现象,且在ReLU回归中达到近最优
- 适用于任意宽度深度,解决广义Barron空间逼近率难题
路径正则化已被证明是训练神经网络的有效正则化方法,其泛化性能优于常见的权重衰减等方法。本文首次为具有路径正则化的多层神经网络提出了一个近乎完整的非渐近泛化理论,适用于一般学习问题。该理论不依赖于损失函数有界这一常见假设。我们的分析超越了偏差-方差权衡框架,与深度学习中的典型现象一致,显著区别于已有结果。具体而言,我们为满足σ(0)=0且损失函数为足够宽的Lipschitz函数的多层网络,给出了显式的泛化误差上界,无需网络宽度、深度或超参数趋于无穷,也无需特定架构(如稀疏性、范数有界)、特定优化算法或损失有界性假设,同时考虑了近似误差。理论的关键特征在于纳入了近似误差。特别地,我们解决了Weinan E等人提出的关于广义Barron空间中逼近率的开放问题。此外,我们证明了该理论在带ReLU激活的回归问题中具有近最小最大最优性。值得注意的是,我们的上界明确表现出著名的双下降现象,这是与现有结果最显著的区别。我们认为,该理论很可能揭示了双下降现象的真实机制。
原文摘要 · Abstract (English)
Path regularization has shown to be a very effective regularization to train neural networks, leading to a better generalization property than common regularizations i.e. weight decay, etc. We propose a first near-complete (as will be made explicit in the main text) nonasymptotic generalization theory for multilayer neural networks with path regularizations for general learning problems. In particular, it does not require the boundedness of the loss function, as is commonly assumed in the literature. Our theory goes beyond the bias-variance tradeoff and aligns with phenomena typically encountered in deep learning. It is therefore sharply different from other existing nonasymptotic generalization error bounds. More explicitly, we propose an explicit generalization error upper bound for multilayer neural networks with $σ(0)=0$ and sufficiently broad Lipschitz loss functions, without requiring the width, depth, or other hyperparameters of the neural network to approach infinity, a specific neural network architecture (e.g., sparsity, boundedness of some norms), a particular optimization algorithm, or boundedness of the loss function, while also taking approximation error into consideration. A key feature of our theory is that it also considers approximation errors. In particular, we solve an open problem proposed by Weinan E et. al. regarding the approximation rates in generalized Barron spaces. Furthermore, we show the near-minimax optimality of our theory for regression problems with ReLU activations. Notably, our upper bound exhibits the famous double descent phenomenon for such networks, which is the most distinguished characteristic compared with other existing results. We argue that it is highly possible that our theory reveals the true underlying mechanism of the double descent phenomenon.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。