arXiv:2605.16622cs.LGmath.OC2026-05被引 1

权重衰减通过抑制参数空间中的锐化过程增强训练稳定性。

Does Weight Decay Enhance Training Stability?

论文配图:Does Weight Decay Enhance Training Stability?
图 1 · 摘自论文原文
  • 分析权重衰减在边缘稳定区的动态,发现其能有效减缓渐进式锐化。
  • 在CNN中抑制振荡,在MLP中引发相变,锐度稳定在理论值以下。
  • 揭示参数向量与锐度梯度全局对齐是相变的机制,适用于函数空间搜索稳定。

在现代深度学习中,权重衰减常被归功于‘稳定’训练动态,与其经典的静态正则化角色不同。本文探讨核心问题:权重衰减是否真的稳定训练?若然,机制为何?我们从参数空间动态和损失锐度角度出发,分析其在边缘稳定(EoS)下的影响。结果表明,权重衰减能稳健地减缓‘渐进式锐化’。此外,我们发现显著的架构依赖性相变:在卷积网络中,权重衰减抑制了边缘稳定区的振荡;而在多层感知机中,增加权重衰减会引发相变,使锐度在远低于理论边界 $\frac{2}{\eta}$ 处稳定。我们构建了一个数学框架准确建模这些现象,并识别出参数向量与锐度梯度的全局对齐是相变的机制驱动。更重要的是,这些现象转化为函数空间(NTK)中的搜索稳定性。最后,我们指出基于凸/二次启发的曲率阈值在正则化下可能不可靠作为稳定性诊断依据。

原文摘要 · Abstract (English)

In modern deep learning, weight decay is often credited with "stabilizing" training dynamics, diverging from its classical role as a static regularization penalty. We investigate a fundamental question: *does weight decay stabilize training dynamics, and if so, through which mechanism?* Indeed, training stability is understood through different but related notions in the literature. We consider how weight decay affects the parameter-space dynamics and loss sharpness by analyzing its effects at the \emph{Edge of Stability} (EoS). We show that weight decay robustly slows *progressive sharpening}. Furthermore, we uncover a striking architecture-dependent phase transition. In CNNs, weight decay dampens the oscillations at the EoS, while in MLPs, increasing weight decay causes a phase transition in which the sharpness stabilizes at a threshold significantly below the theoretical $\frac{2}η$ boundary. We develop a mathematical framework that accurately models these phenomena and identify the global alignment of the parameter vector and the sharpness gradient as the mechanistic driver of the phase transition. Importantly, we show that these phenomena translate into stability in terms of search in function-space (NTK). Last, this shows that curvature thresholds obtained from convex/quadratic heuristics may not be reliable stability diagnostics under regularization.

权重衰减训练稳定边缘稳定模型架构

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。