arXiv:2411.01714cs.LG2024-11

发现SAM的泛化优势源于近似计算,而非原设计机制。

1st-Order Magic: Analysis of Sharpness-Aware Minimization

  • 用近似方法优化损失平坦性,提升模型泛化能力
  • 越精确的近似反而降低泛化性能,说明效果依赖近似本身
  • 揭示优化中近似项的关键作用,适合研究优化机制者阅读

Sharpness-Aware Minimization (SAM) 是一种通过偏好更平坦的损失极小值来提升泛化能力的优化技术。为实现此目标,SAM 使用计算高效的近似方法优化一个惩罚尖锐度的改进目标。然而,我们发现对所提 SAM 目标更精确的近似会损害泛化性能,表明 SAM 的泛化优势实际上根植于这些近似,而非原始设计机制。这揭示了对 SAM 有效性理解的缺口,并呼吁进一步研究近似在优化中的作用。

原文摘要 · Abstract (English)

Sharpness-Aware Minimization (SAM) is an optimization technique designed to improve generalization by favoring flatter loss minima. To achieve this, SAM optimizes a modified objective that penalizes sharpness, using computationally efficient approximations. Interestingly, we find that more precise approximations of the proposed SAM objective degrade generalization performance, suggesting that the generalization benefits of SAM are rooted in these approximations rather than in the original intended mechanism. This highlights a gap in our understanding of SAM's effectiveness and calls for further investigation into the role of approximations in optimization.

优化器泛化近似分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。