无需修改模型即可获得深度网络的非平凡泛化界
Non-vacuous Generalization Bounds for Deep Neural Networks without any modification to the trained models
- 基于数据分布设计可精确计算的泛化边界
- 在包含6亿参数的ImageNet模型上仍保持非平凡界限
- 揭示数据几何与模型局部行为共同决定泛化性能
理解并验证现代深度神经网络的行为仍是可靠机器学习中的根本挑战。我们提出一类直接适用于训练后模型的数据相关泛化边界,无需任何修改。特别地,我们给出了一个可精确计算的边界,在所有评估过的网络中均保持非平凡性,包括具有6亿参数的ImageNet规模模型。这是首次证明即使对大型、未经修改的深度网络也能获得有意义的泛化保证。我们的方法揭示了泛化由训练模型与数据分布几何之间的相互作用所决定。我们将泛化误差分解为两个可解释成分:一个分布复杂度项,刻画数据质量在输入空间中的分布;一个局部模型行为项,刻画网络在各个区域内的表现。这种联合依赖关系指出了泛化差距出现的位置与原因。实验表明,边界的部分成分对真实测试误差有很强预测力,且当划分与数据内在几何一致时,边界更紧,凸显数据依赖的局部正则性是泛化的重要驱动力。
原文摘要 · Abstract (English)
Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning. We introduce a new class of data-dependent generalization bounds that apply directly to trained models, without any modification. In particular, we present an exactly computable bound that is non-vacuous across all evaluated networks, including ImageNet-scale models with 600M parameters. This this is the first work showing that meaningful generalization guarantees are achievable even for large, unaltered deep networks. Our approach reveals that generalization is governed by the interaction between the trained model and the geometry of the data distribution. We decompose the generalization error into two interpretable components: a distributional complexity term, capturing how the data mass is distributed across the input space, and local model-behavior terms, capturing the network's behavior within individual regions. This joint dependence identifies where and why generalization gaps arise. Empirically, some components of our bound are highly predictive of the true test error, and the bound tightens when the partition aligns with the intrinsic data geometry, highlighting data-dependent local regularity as a key driver of generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。