解释过参数模型为何能泛化良好,发现数据越多坏解越少
Rethinking generalization of classifiers in separable classes scenarios and over-parameterized regimes
- 分析可分数据下分类器学习过程,揭示好解比例随数据指数上升
- 证明泛化性能与模型复杂度无关,仅取决于真实误差分布密度
- 为神经网络超参数下的优异泛化提供理论解释,适合研究者参考
我们研究了在类别可分或分类器过参数化场景下的学习动态。在这两种情况下,经验风险最小化(ERM)都会导致零训练误差。然而,存在多个全局最小值,其中一些泛化良好,一些则不然。我们发现,在类别可分场景中,'坏'全局最小值的比例随训练样本数n呈指数下降。我们的分析给出了仅依赖于给定分类器函数集的真实误差密度分布的边界和学习曲线,与该集合的大小或复杂度(如参数数量)无关。这一发现可能揭示了过参数化神经网络出人意料的良好泛化能力。对于过参数化情形,我们提出一个真实误差密度分布模型,所得学习曲线与在MNIST和CIFAR-10上的实验结果一致。
原文摘要 · Abstract (English)
We investigate the learning dynamics of classifiers in scenarios where classes are separable or classifiers are over-parameterized. In both cases, Empirical Risk Minimization (ERM) results in zero training error. However, there are many global minima with a training error of zero, some of which generalize well and some of which do not. We show that in separable classes scenarios the proportion of "bad" global minima diminishes exponentially with the number of training data n. Our analysis provides bounds and learning curves dependent solely on the density distribution of the true error for the given classifier function set, irrespective of the set's size or complexity (e.g., number of parameters). This observation may shed light on the unexpectedly good generalization of over-parameterized Neural Networks. For the over-parameterized scenario, we propose a model for the density distribution of the true error, yielding learning curves that align with experiments on MNIST and CIFAR-10.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。