arXiv:2412.05169cs.LGcs.AI2024-12被引 1

SAM算法在分布外泛化中表现优异,优于传统优化器。

Towards Understanding the Role of Sharpness-Aware Minimization Algorithms for Out-of-Distribution Generalization

  • 采用尖锐度感知最小化提升模型对未知分布的适应能力。
  • 在零样本分布外场景下,最优SAM变体比Adam提升8.01%。
  • 适用于需要强泛化能力的跨域任务,如领域自适应与迁移学习。

近期,尖锐度感知最小化(SAM)因其在降低尖锐度以提升泛化能力方面的潜力而受到关注。本文系统评估了八种SAM变体在零样本分布外(OOD)泛化中的表现,发现原始SAM比Adam基准提升4.76%,最强变体平均提升8.01%。研究还给出了该设置下的分布外泛化界。进一步地,在渐进域自适应(GDA)场景中,原始SAM在各数据集上平均优于Adam 0.82%,最强变体平均提升1.52%。相应地,提出了GDA场景下的泛化界,但其渐近性能并不优于现有自训练方法,表明当前理论尚无法完全解释SAM的实证优势。未来工作可探索更紧致的分析框架。

原文摘要 · Abstract (English)

Recently, sharpness-aware minimization (SAM) has emerged as a promising method to improve generalization by minimizing sharpness, which is known to correlate well with generalization ability. Since the original proposal of SAM, many variants of SAM have been proposed to improve its accuracy and efficiency, but comparisons have mainly been restricted to the i.i.d. setting. In this paper we study SAM for out-of-distribution (OOD) generalization. First, we perform a comprehensive comparison of eight SAM variants on zero-shot OOD generalization, finding that the original SAM outperforms the Adam baseline by $4.76\%$ and the strongest SAM variants outperform the Adam baseline by $8.01\%$ on average. We then provide an OOD generalization bound in terms of sharpness for this setting. Next, we extend our study of SAM to the related setting of gradual domain adaptation (GDA), another form of OOD generalization where intermediate domains are constructed between the source and target domains, and iterative self-training is done on intermediate domains, to improve the overall target domain error. In this setting, our experimental results demonstrate that the original SAM outperforms the baseline of Adam on each of the experimental datasets by $0.82\%$ on average and the strongest SAM variants outperform Adam by $1.52\%$ on average. We then provide a generalization bound for SAM in the GDA setting. Asymptotically, this generalization bound is no better than the one for self-training in the literature of GDA. This highlights a further disconnection between the theoretical justification for SAM versus its empirical performance, with recent work finding that low sharpness alone does not account for all of SAM's generalization benefits. For future work, we provide several potential avenues for obtaining a tighter analysis for SAM in the OOD setting.

分布外泛化尖锐度感知模型泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。