让优化器主动找平坦的解,提升模型泛化与不确定性估计。
Flatness-Aware Stochastic Gradient Langevin Dynamics
- 通过噪声与温度耦合机制,引导优化过程偏向平坦区域。
- 理论证明可收敛到偏平坦的吉布斯分布,且有明确风险保证。
- 适用于图像分类、异常检测等任务,结果稳定可信。
损失曲面的平坦性被广泛认为是理解深度学习算法行为与泛化能力的重要视角。受此启发,我们提出一种一阶优化方法——平坦性感知随机梯度朗之万动力学(fSGLD),在保持SGD和SGLD计算与内存效率的同时,引导学习过程向平坦的极小值区域偏移。我们提供了非渐近理论分析,表明当噪声尺度σ与逆温度β按理论指定方式耦合时,fSGLD可收敛至一个偏平坦的吉布斯分布,并给出明确的超出风险保证。我们在标准优化器基准、贝叶斯图像分类、不确定性量化及分布外检测任务上进行实证评估,结果表明fSGLD性能稳定且不确定性估计可靠。额外实验验证了理论规定的β-σ耦合优于解耦设置。
原文摘要 · Abstract (English)
Flatness of the loss landscape has been widely studied as an important perspective for understanding the behavior and generalization of deep learning algorithms. Motivated by this view, we propose Flatness-Aware Stochastic Gradient Langevin Dynamics (fSGLD), a first-order optimization method that biases learning its dynamics toward flat basins while retaining the computational and memory efficiency of SGD and SGLD. We provide a non-asymptotic theoretical analysis showing that fSGLD targets a flatness-biased Gibbs distribution under a theoretically prescribed coupling between the noise scale $σ$ and the inverse temperature $β$, together with explicit excess risk guarantees. We empirically evaluate fSGLD across standard optimizer benchmarks, Bayesian image classification, uncertainty quantification, and out-of-distribution detection, demonstrating consistently strong performance and reliable uncertainty estimates. Additional experiments confirm the effectiveness of the theoretically prescribed $β$-$σ$ coupling compared to decoupled choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。