证明了有限宽度ReLU网络在任意回归函数下存在良性过拟合。
Benign Overfitting for Regression with Trained Two-Layer ReLU Networks
- 用梯度流训练两层ReLU网络,不假设回归函数和噪声形式。
- 在神经正切核范围内,模型过拟合并保持良好泛化性能。
- 为非理想数据下的深度学习提供了理论支持,适合理论研究者。
我们研究使用两层全连接神经网络(带ReLU激活)并以梯度流训练的最小二乘回归问题。首个结果是无需对潜在回归函数或噪声做任何假设(仅要求有界),即可获得泛化误差界。在神经正切核(Neural Tangent Kernel, NTK)范式下,通过将超出风险分解为估计误差与近似误差,将梯度流视为隐式正则化器。该分解视角为神经网络中的梯度下降提供了新理解,避免了统一收敛的陷阱。本工作还证明,在相同设定下,训练后的网络会过拟合数据。上述结果首次建立了针对任意回归函数的有限宽度ReLU网络的良性过拟合理论。
原文摘要 · Abstract (English)
We study the least-square regression problem with a two-layer fully-connected neural network, with ReLU activation function, trained by gradient flow. Our first result is a generalization result, that requires no assumptions on the underlying regression function or the noise other than that they are bounded. We operate in the neural tangent kernel regime, and our generalization result is developed via a decomposition of the excess risk into estimation and approximation errors, viewing gradient flow as an implicit regularizer. This decomposition in the context of neural networks is a novel perspective of gradient descent, and helps us avoid uniform convergence traps. In this work, we also establish that under the same setting, the trained network overfits to the data. Together, these results, establishes the first result on benign overfitting for finite-width ReLU networks for arbitrary regression functions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。