arXiv:2601.01465cs.LG2026-01ICLR被引 2

改进了基于信息论的泛化界,让梯度下降更懂平坦性。

Leveraging Flatness to Improve Information-Theoretic Generalization Bounds for SGD

  • 用‘全知轨迹’技术增强对优化路径平坦性的捕捉
  • 实验显示新界比旧界紧40%以上,且与平坦性正相关
  • 适合研究泛化理论、模型优化的学者参考

信息论泛化界被用于分析学习算法的泛化能力,其固有地依赖数据和算法特性,因此可利用数据与算法属性获得更紧的界。然而我们发现,尽管平坦性偏差对SGD泛化至关重要,现有信息论界未能捕捉到更好平坦性带来的泛化提升,且数值上仍较松散。这源于现有界对SGD平坦性偏好的利用不足。本文推导出一种更注重平坦性的信息论界,表明最终权重协方差中高方差方向在损失曲面中局部曲率越小,模型泛化越好。在深度神经网络上的实验表明,该界不仅能正确反映平坦性提升时的更好泛化,且数值上显著更紧。此效果通过‘全知轨迹’灵活技术实现。应用于凸-利普希茨-有界问题下的梯度下降最小最大过拟合风险,将代表性信息论界的Ω(1)率改进为O(1/√n),还暗示可绕过记忆-泛化权衡。

原文摘要 · Abstract (English)

Information-theoretic (IT) generalization bounds have been used to study the generalization of learning algorithms. These bounds are intrinsically data- and algorithm-dependent so that one can exploit the properties of data and algorithm to derive tighter bounds. However, we observe that although the flatness bias is crucial for SGD's generalization, these bounds fail to capture the improved generalization under better flatness and are also numerically loose. This is caused by the inadequate leverage of SGD's flatness bias in existing IT bounds. This paper derives a more flatness-leveraging IT bound for the flatness-favoring SGD. The bound indicates the learned models generalize better if the large-variance directions of the final weight covariance have small local curvatures in the loss landscape. Experiments on deep neural networks show our bound not only correctly reflects the better generalization when flatness is improved, but is also numerically much tighter. This is achieved by a flexible technique called "omniscient trajectory". When applied to Gradient Descent's minimax excess risk on convex-Lipschitz-Bounded problems, it improves representative IT bounds' $Ω(1)$ rates to $O(1/\sqrt{n})$. It also implies a by-pass of memorization-generalization trade-offs.

泛化理论SGD信息论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。