提出统一框架,证明SAM在非凸问题中更优收敛性。
Sharpness-Aware Minimization: General Analysis and Improved Rates
- 设计统一更新规则,同时涵盖SAM与USAM
- 在弱噪声假设下实现非凸优化的收敛保证
- 支持任意采样策略,适合深度学习训练
Sharpness-Aware Minimization (SAM) 通过最小化损失曲面的尖锐度,显著提升模型泛化能力。然而,其在非凸设置下的收敛性质仍存在若干未解问题,包括更新规则中归一化的收益、对严格有界方差假设的依赖,以及不同采样策略下的收敛保证。本文提出统一的SAM(Unified SAM)更新规则,统一分析SAM及其无归一化变体(USAM),并在更宽松自然的随机噪声假设下给出收敛结果。理论表明,在满足Polyak-Lojasiewicz(PL)条件的非凸函数上,该算法在不同步长选择下均具备收敛性。所提理论适用于任意采样范式(包含重要性采样为特例),可分析文献中未明确考虑的SAM变体。实验验证了理论结论,并进一步展示了Unified SAM在图像分类任务中训练深层神经网络的实用性。
原文摘要 · Abstract (English)
Sharpness-Aware Minimization (SAM) has emerged as a powerful method for improving generalization in machine learning models by minimizing the sharpness of the loss landscape. However, despite its success, several important questions regarding the convergence properties of SAM in non-convex settings are still open, including the benefits of using normalization in the update rule, the dependence of the analysis on the restrictive bounded variance assumption, and the convergence guarantees under different sampling strategies. To address these questions, in this paper, we provide a unified analysis of SAM and its unnormalized variant (USAM) under one single flexible update rule (Unified SAM), and we present convergence results of the new algorithm under a relaxed and more natural assumption on the stochastic noise. Our analysis provides convergence guarantees for SAM under different step size selections for non-convex problems and functions that satisfy the Polyak-Lojasiewicz (PL) condition (a non-convex generalization of strongly convex functions). The proposed theory holds under the arbitrary sampling paradigm, which includes importance sampling as special case, allowing us to analyze variants of SAM that were never explicitly considered in the literature. Experiments validate the theoretical findings and further demonstrate the practical effectiveness of Unified SAM in training deep neural networks for image classification tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。