arXiv:2502.08448cs.LGstat.ML2025-02被引 1

提出新型优化方法M-SAM,让模型训练更稳定、泛化更好。

Monge SAM: Robust Reparameterization-Invariant Sharpness-Aware Minimization Based on Loss Geometry

  • 基于损失曲面的黎曼度量设计新优化器,避免参数重参数影响
  • 在多模态对齐任务中表现优于SAM,收敛更鲁棒
  • 无需复杂假设,计算开销与SAM相当,适合实际部署

深度神经网络的研究表明,损失函数曲面的平坦极小值与更好的泛化能力相关。尖锐感知最小化(SAM)通过在对抗扰动处计算梯度来寻找平坦区域,但其扰动依赖欧氏度量,不满足重参数化不变性,导致尖锐性与泛化关系模糊。本文提出基于损失曲面自然诱导的黎曼度量的蒙日SAM(M-SAM),实现重参数化不变的尖锐感知最小化。相比以往方法,M-SAM适用于任意建模选择,仅需温和假设且计算效率与SAM相当。理论分析表明,M-SAM介于SAM与梯度下降之间,增强对超参数的鲁棒性,并减少对鞍点等次优平衡点的吸引。我们在多模态表示对齐任务中同时从理论上和实证上验证了该特性。

原文摘要 · Abstract (English)

Recent studies on deep neural networks show that flat minima of the loss landscape correlate with improved generalization. Sharpness-aware minimization (SAM) efficiently finds flat regions by updating the parameters according to the gradient at an adversarial perturbation. The perturbation depends on the Euclidean metric, making SAM non-invariant under reparametrizations, which blurs sharpness and generalization. We propose Monge SAM (M-SAM), a reparametrization invariant version of SAM by considering a Riemannian metric in the parameter space induced naturally by the loss surface. Compared to previous approaches, M-SAM works under any modeling choice, relies only on mild assumptions while being as computationally efficient as SAM. We theoretically argue that M-SAM varies between SAM and gradient descent (GD), which increases robustness to hyperparameter selection and reduces attraction to suboptimal equilibria like saddle points. We demonstrate this behavior both theoretically and empirically on a multi-modal representation alignment task.

优化算法泛化提升黎曼几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。