提出更抗噪的加速梯度方法,提升深度学习训练稳定性
SHANG++: Robust Stochastic Acceleration under Multiplicative Noise
- 基于哈密顿驱动的加速流,设计半隐式离散算法
- 在乘性噪声下收敛更快,单配置保持误差≤1%
- 适合对噪声敏感的深度学习任务,参数调优少
在乘性噪声缩放(MNS)条件下,原始Nesterov加速方法对噪声敏感,可能发散。本文通过离散海森驱动的Nesterov加速梯度流,提出两种加速随机梯度下降方法。首先推导出SHANG,一种直接的半隐式离散化,已提升MNS下的稳定性。随后引入SHANG++,加入阻尼修正项,实现更快收敛与更强噪声鲁棒性。我们在凸和强凸目标下建立了收敛保证,并给出明确参数选择。实验表明,SHANG++在各类凸问题及深度学习应用中表现稳定;在ResNet-34的专用噪声实验中,单一超参数配置使准确率仅比无噪声设置低1个百分点。所有实验中,SHANG++在鲁棒性和效率上均优于现有加速方法,且对参数变化不敏感。
原文摘要 · Abstract (English)
Under the multiplicative noise scaling (MNS) condition, original Nesterov acceleration is provably sensitive to noise and may diverge when gradient noise overwhelms the signal. In this paper, we develop two accelerated stochastic gradient descent methods by discretizing the Hessian-driven Nesterov accelerated gradient flow. We first derive SHANG, a direct semi-implicit discretization that already improves stability under MNS. We then introduce SHANG++, which adds a damping correction and achieves faster convergence with greater noise robustness. We establish convergence guarantees for both convex and strongly convex objectives under MNS, together with explicit parameter choices. In our experiments, SHANG++ performs consistently well across convex problems and applications in deep learning. In a dedicated noise experiment on ResNet-34, a single hyperparameter configuration maintains accuracy within one percentage point of the noise-free setting. Across all experiments, SHANG++ outperforms existing accelerated methods in robustness and efficiency, with minimal parameter sensitivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。