arXiv:2505.11972cs.LGstat.ML2025-05被引 1

通过分离梯度方向与曲率主空间,加速训练并提升稳定性。

Accelerating Neural Network Training Along Sharp and Flat Directions

  • 限制更新在主曲率空间的正交方向,聚焦平坦区域优化。
  • 平坦方向更新可加速收敛,但需调参避免不稳定。
  • 提出统一优化器,融合主流方法,适合追求高效训练的研究者。

近期研究发现,神经网络训练中梯度与海森矩阵主特征空间(主导子空间)高度对齐。与此同时,人们对海森谱中陡峭与平坦方向的不同作用日益关注。本文研究了批量梯度下降(Bulk-SGD),其将更新限制在主导子空间的正交补空间。通过消融实验,我们刻画了Bulk-SGD的稳定性特性,并识别出决定其行为的关键超参数。结果表明,沿批量子空间(对应损失景观中的平坦方向)的更新可加速收敛,但可能损害稳定性。为平衡此矛盾,我们提出一种插值梯度方法,统一了SGD、Dom-SGD和Bulk-SGD。最后,我们实证连接该子空间分解与广义高斯-牛顿项及函数海森项,发现曲率能量主要集中在主导子空间。研究揭示了设计曲率感知优化器的系统性思路。

原文摘要 · Abstract (English)

Recent work has highlighted a surprising alignment between gradients and the top eigenspace of the Hessian -- termed the Dominant subspace -- during neural network training. Concurrently, there has been growing interest in the distinct roles of sharp and flat directions in the Hessian spectrum. In this work, we study Bulk-SGD, a variant of SGD that restricts updates to the orthogonal complement of the Dominant subspace. Through ablation studies, we characterize the stability properties of Bulk-SGD and identify critical hyperparameters that govern its behavior. We show that updates along the Bulk subspace, corresponding to flatter directions in the loss landscape, can accelerate convergence but may compromise stability. To balance these effects, we introduce interpolated gradient methods that unify SGD, Dom-SGD, and Bulk-SGD. Finally, we empirically connect this subspace decomposition to the Generalized Gauss-Newton and Functional Hessian terms, showing that curvature energy is largely concentrated in the Dominant subspace. Our findings suggest a principled approach to designing curvature-aware optimizers.

优化器海森矩阵训练加速梯度分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。