arXiv:2503.04111cs.LGcs.AI2025-03

揭示过参数神经网络为何能泛化,关键在表达能力与数据量匹配。

Generalizability of Neural Networks Minimizing Empirical Risk Based on Expressive Ability

  • 基于网络表达能力推导泛化下界,不依赖严格假设。
  • 证明需足够多样本才能泛化,且样本数常超网络容量。
  • 解释深度学习中过参数化、鲁棒性等现象的理论根源。

学习方法的核心目标是泛化。经典统一泛化界依赖VC维或Rademacher复杂度,无法解释深度学习中过参数模型表现出的良好泛化能力。而算法相关泛化界(如稳定性界)常依赖严苛假设。本文研究最小化或近似最小化经验风险的神经网络的泛化能力,建立基于网络表达能力的泛化下界:当训练样本和网络规模足够大时,包括过参数化模型在内的网络可有效泛化。此外,我们给出泛化的必要条件,表明在某些数据分布下,确保泛化的样本量需超过表示该分布所需的网络规模。最后,为深度学习中的鲁棒泛化、过参数化的重要性及损失函数对泛化的影响提供理论洞察。

原文摘要 · Abstract (English)

The primary objective of learning methods is generalization. Classic uniform generalization bounds, which rely on VC-dimension or Rademacher complexity, fail to explain the significant attribute that over-parameterized models in deep learning exhibit nice generalizability. On the other hand, algorithm-dependent generalization bounds, like stability bounds, often rely on strict assumptions. To establish generalizability under less stringent assumptions, this paper investigates the generalizability of neural networks that minimize or approximately minimize empirical risk. We establish a lower bound for population accuracy based on the expressiveness of these networks, which indicates that with an adequate large number of training samples and network sizes, these networks, including over-parameterized ones, can generalize effectively. Additionally, we provide a necessary condition for generalization, demonstrating that, for certain data distributions, the quantity of training data required to ensure generalization exceeds the network size needed to represent the corresponding data distribution. Finally, we provide theoretical insights into several phenomena in deep learning, including robust generalization, importance of over-parameterization, and effect of loss function on generalization.

泛化能力神经网络过参数化理论分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。