提出与架构无关的深度ReLU网络泛化界,突破过参数化限制。
Architecture independent generalization bounds for overparametrized deep ReLU networks
- 基于度量几何与权重范数构建泛化上界,不依赖VC维和网络结构。
- 在样本数不超过输入维时,可显式构造零损失解。
- 理论预测与MNIST实验结果偏差仅22%,适用于高维学习场景。
我们证明了过参数化神经网络的测试误差可独立于过参数化程度和Vapnik-Chervonenkis(VC)维。给出的显式上界仅依赖于测试集与训练集的度量几何、激活函数的正则性,以及权重的算子范数和偏置的范数。对于训练样本数不超过输入空间维度的过参数化深层ReLU网络,我们显式构造出无需梯度下降的零损失最小化器,并证明了与网络架构无关的一致泛化界。在MNIST上的计算实验表明,理论预测与真实测试误差平均偏差在22%以内。
原文摘要 · Abstract (English)
We prove that overparametrized neural networks are able to generalize with a test error that is independent of the level of overparametrization, and independent of the Vapnik-Chervonenkis (VC) dimension. We prove explicit bounds that only depend on the metric geometry of the test and training sets, on the regularity properties of the activation function, and on the operator norms of the weights and norms of biases. For overparametrized deep ReLU networks with a training sample size bounded by the input space dimension, we explicitly construct zero loss minimizers without use of gradient descent, and prove a uniform generalization bound that is independent of the network architecture. We perform computational experiments of our theoretical results with MNIST, and obtain agreement with the true test error within a 22 % margin on average.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。