arXiv:2604.00505cs.LGcs.AI2026-04

提出新方法,让过参数化神经网络的泛化界不再空洞。

Towards Initialization-dependent and Non-vacuous Generalization Bounds for Overparameterized Shallow Neural Networks

论文配图:Towards Initialization-dependent and Non-vacuous Generalization Bounds for Overparameterized Shallow Neural Networks
图 1 · 摘自论文原文
  • 用路径范数衡量初始点距离,改进传统方法
  • 理论证明可得非空洞的泛化误差上界
  • 适合研究模型泛化性与初始化关系的研究者

过参数化神经网络在参数量超过训练样本数时仍表现出优异的泛化能力,这种现象称为良性过拟合。现有分析多基于从初始化出发的距离的Frobenius范数,但在过参数化场景下常导致空洞的泛化界。本文针对具有广义Lipschitz激活函数的浅层神经网络,提出基于路径范数的初始化依赖复杂度界,通过引入新的剥皮技术处理初始化约束带来的挑战,并给出紧致下界(相差常数因子)。实验表明,该分析可导出对过参数化网络有效的非空洞泛化界。

原文摘要 · Abstract (English)

Overparameterized neural networks often show a benign overfitting property in the sense of achieving excellent generalization behavior despite the number of parameters exceeding the number of training examples. A promising direction to explain benign overfitting is to relate generalization to the norm of distance from initialization, motivated by the empirical observations that this distance is often significantly smaller than the norm itself. However, the existing initialization-dependent complexity analyses measure the distance from initialization by the Frobenius norm, and often imply vacuous bounds in practice for overparamterized models. In this paper, we develop initialization-dependent complexity bounds for shallow neural networks with general Lipschitz activation functions. Our bounds depend on the path-norm of the distance from initialization, which are derived by introducing a new peeling technique to handle the challenge along with the initialization-dependent constraint. We also develop a lower bound tight up to a constant factor. Finally, we conduct empirical comparisons and show that our generalization analysis implies non-vacuous bounds for overparameterized networks.

泛化界神经网络初始化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。