提出新泛化界,让深度模型误差估计更紧实、更可用。
Upper Bounds on the Generalization Error of Deep Learning Models via Local Robustness and Stability

- 按输入空间子区域划分,动态调整鲁棒性权重
- 在ImageNet上估计误差紧致,优于已有方法
- 适合关注模型可靠性与真实性能评估的研究者
泛化能力是数据驱动模型(尤其是安全关键场景下的深度学习模型)的关键属性。基于鲁棒性的泛化界因其能将鲁棒性与泛化性能关联而受到关注,通常是数据相关的。然而,现有边界在实际中常因过于宽松而失效,其上界远高于实际误差,限制了真实评估的实用性。问题不仅源于不确定性项,更在于鲁棒性项本身,尤其对0-1损失而言。现有方法通常将鲁棒性视为全局度量,忽略了输入空间各子区域间的差异。本文提出一种新泛化界,通过根据每个子区域内稳定与不稳定样本数量来缩放鲁棒性项,以克服这一缺陷。该边界同时包含数据和模型相关因素,保持实际应用价值(得到更紧的真误差上界)。在ImageNet训练的模型上实验表明,本方法始终非空洞,且在各类鲁棒深度神经网络中,估计结果最为紧密,与实际表现高度一致。
原文摘要 · Abstract (English)
Generalization is a critical property of data-driven models, particularly deep learning models deployed in safety-critical applications. Robustness-based generalization bounds have gained attention as a principled way to link robustness properties to generalization performance, often in a data-dependent manner. However, most existing bounds suffer from vacuousness in practical settings, yielding loose upper bounds that greatly exceed the actual error rates and limiting their usefulness for real-world evaluation. While this issue is often attributed to the uncertainty term, a substantial part of the problem originates from the robustness term itself, particularly for the 0-1 loss. Existing approaches typically treat the robustness term as a global measure, ignoring its variation across different sub-regions of the input space. In this work, we propose a generalization bound that addresses this limitation by scaling the robustness term according to the number of stable and unstable samples within each sub-region. Our bounds incorporate both data- and model-dependent factors while maintaining practical relevance (yielding tighter upper bounds on true error). Experiments on models trained on the ImageNet dataset show that our bounds remain consistently non-vacuous and achieve the tightest estimates among existing methods, closely aligning with empirical performance across a range of robust deep neural networks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。