arXiv:2604.02653cs.LG2026-04被引 1

证明了特定损失函数下梯度下降在稳定边缘仍能收敛

Product-Stability: Provable Convergence for Gradient Descent on the Edge of Stability

  • 提出'乘积稳定性'结构特性,刻画损失函数的收敛性质
  • 在稳定边缘训练时,模型仍可收敛至局部最小值
  • 适用于交叉熵等常见损失,适合研究训练动态的学者

实践中,现代深度学习训练常处于稳定边缘(EoS),此时损失函数的尖锐度超过经典收敛理论的阈值。尽管已有进展,现有理论解释或依赖严格假设,或仅针对特定平方损失目标。本文引入并研究损失函数的一种结构性质——乘积稳定性。我们证明:对于具有乘积稳定极小值的损失函数,当目标形式为 $(x,y) o l(xy)$ 时,梯度下降在 EoS 区域仍可保证收敛至局部最小值。该框架显著推广了先前结果,适用于广泛损失类型,包括二分类交叉熵。通过分岔图刻画训练动态,解释稳定振荡的出现,并精确量化收敛时的尖锐度。结果为更广泛的损失函数提供了稳定 EoS 训练的理论依据。

原文摘要 · Abstract (English)

Empirically, modern deep learning training often occurs at the Edge of Stability (EoS), where the sharpness of the loss exceeds the threshold below which classical convergence analysis applies. Despite recent progress, existing theoretical explanations of EoS either rely on restrictive assumptions or focus on specific squared-loss-type objectives. In this work, we introduce and study a structural property of loss functions that we term product-stability. We show that for losses with product-stable minima, gradient descent applied to objectives of the form $(x,y) \mapsto l(xy)$ can provably converge to the local minimum even when training in the EoS regime. This framework substantially generalizes prior results and applies to a broad class of losses, including binary cross entropy. Using bifurcation diagrams, we characterize the resulting training dynamics, explain the emergence of stable oscillations, and precisely quantify the sharpness at convergence. Together, our results offer a principled explanation for stable EoS training for a wider class of loss functions.

优化理论稳定边缘梯度下降损失函数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。