arXiv:2510.18934cs.LG2025-10

多数深度学习泛化度量对训练微调敏感,结果不可靠。

Position: Many generalization measures for deep learning are fragile

  • 检验训练后模型的泛化度量易受超参数微调影响
  • 学习率调整可使路径范数等度量的曲线斜率反转
  • 建议新度量需主动评估其脆弱性,尤其关注数据复杂度

本文指出,许多训练后计算的泛化度量存在脆弱性:对深层神经网络进行微小训练调整(如学习率变化或优化器切换),虽几乎不影响模型性能,却可能显著改变度量值、趋势或缩放行为。例如,细微的超参数变化即可导致路径范数等常用度量的学习曲线斜率反转。此外,尽管基于PAC-Bayes的起源度量通常被认为可靠,对超参数变化不敏感,但其无法捕捉学习曲线中数据复杂度的差异。相较之下,函数基的边际似然型PAC-Bayes界能反映数据复杂度及其缩放行为,但不属于训练后度量。本文主张,新泛化度量的开发者应主动评估其脆弱性。

原文摘要 · Abstract (English)

In this position paper, we argue that many post-mortem generalization measures -- those computed on trained networks -- are \textbf{fragile}: small training modifications that barely affect the performance of the underlying deep neural network can substantially change a measure's value, trend, or scaling behavior. For example, minor hyperparameter changes, such as learning rate adjustments or switching between SGD variants, can reverse the slope of a learning curve in widely used generalization measures such as the path norm. We also identify subtler forms of fragility. For instance, the PAC-Bayes origin measure is regarded as one of the most reliable, and is indeed less sensitive to hyperparameter tweaks than many other measures. However, it completely fails to capture differences in data complexity across learning curves. This data fragility contrasts with the function-based marginal-likelihood PAC-Bayes bound, which does capture differences in data-complexity, including scaling behavior, in learning curves, but which is not a post-mortem measure. Beyond demonstrating that many post-mortem bounds are fragile, this position paper also argues that developers of new measures should explicitly audit them for fragility.

泛化度量脆弱性深度学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。