揭示了张量模型中SAM的隐式正则机制,提出更高效的改进方法DAS。
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
- 通过尺度不变性分析核心参数范数动态,发现其受范数与梯度大小协方差调控。
- 提出DAS方法,基于数据自适应缩放核心范数,在多个任务上超越SAM性能。
- 适合关注模型优化、张量分解及高效微调的研究者和工程师。
Sharpness-Aware Minimization (SAM) 被证明是提升过参数化模型泛化能力的有效优化技术。尽管已有研究在简单的双核尺度不变设置中探讨了SAM的隐式正则化行为,但其在更一般的张量化或尺度不变模型中的表现仍不明确。本文利用尺度不变性,分析了一般张量化模型中SAM的范数动态。引入了‘范数偏差’作为核心范数失衡的全局度量,并通过梯度流分析推导出其在SAM下的演化规律。结果表明,SAM对范数偏差的隐式控制由核心范数与其梯度幅度之间的协方差决定。基于此发现,我们提出了一个简单而有效的方法——范数偏差感知缩放(Deviation-Aware Scaling, DAS),通过数据自适应方式显式模拟这一正则化行为。在张量补全、噪声训练、模型压缩和参数高效微调等多个任务上的实验表明,DAS在性能上可媲美甚至超越SAM,同时计算开销更低。
原文摘要 · Abstract (English)
Sharpness-Aware Minimization (SAM) has been proven to be an effective optimization technique for improving generalization in overparameterized models. While prior works have explored the implicit regularization of SAM in simple two-core scale-invariant settings, its behavior in more general tensorized or scale-invariant models remains underexplored. In this work, we leverage scale-invariance to analyze the norm dynamics of SAM in general tensorized models. We introduce the notion of \emph{Norm Deviation} as a global measure of core norm imbalance, and derive its evolution under SAM using gradient flow analysis. We show that SAM's implicit control of Norm Deviation is governed by the covariance between core norms and their gradient magnitudes. Motivated by these findings, we propose a simple yet effective method, \emph{Deviation-Aware Scaling (DAS)}, which explicitly mimics this regularization behavior by scaling core norms in a data-adaptive manner. Our experiments across tensor completion, noisy training, model compression, and parameter-efficient fine-tuning confirm that DAS achieves competitive or improved performance over SAM, while offering reduced computational overhead.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。