解析SGD-M与自适应步长在高维下的行为差异,揭示稳定机制。
High-dimensional limit theorems for SGD: Momentum and Adaptive Step-sizes
- 通过高维极限框架分析SGD-M与自适应步长的动态演化。
- 相同步长下SGD-M会放大高维效应,降低性能;调整后可逼近在线SGD表现。
- 适用于研究高维优化中预条件器如何提升收敛性与稳定性的人群。
我们为带有Polyak动量的随机梯度下降(SGD-M)和自适应步长提出了高维尺度极限。该框架可严格比较在线SGD与其流行变体。我们发现,在适当的时标重缩放和特定步长选择下,SGD-M的极限行为与在线SGD一致;但若步长相同,SGD-M将放大高维效应,可能劣于在线SGD。我们在两个典型学习问题中验证该框架:带尖峰张量的PCA与单指标模型。在高维情形下,基于归一化梯度的自适应步长在线SGD表现出多重优势:其动态具有更接近总体最小值的不动点,并扩大了迭代收敛至解的步长可接受范围。这些结果为早期预条件器在在线SGD失效场景中的稳定与改进作用提供了严格理论支持,契合实证动机。
原文摘要 · Abstract (English)
We develop a high-dimensional scaling limit for Stochastic Gradient Descent with Polyak Momentum (SGD-M) and adaptive step-sizes. This provides a framework to rigourously compare online SGD with some of its popular variants. We show that the scaling limits of SGD-M coincide with those of online SGD after an appropriate time rescaling and a specific choice of step-size. However, if the step-size is kept the same between the two algorithms, SGD-M will amplify high-dimensional effects, potentially degrading performance relative to online SGD. We demonstrate our framework on two popular learning problems: Spiked Tensor PCA and Single Index Models. In both cases, we also examine online SGD with an adaptive step-size based on normalized gradients. In the high-dimensional regime, this algorithm yields multiple benefits: its dynamics admit fixed points closer to the population minimum and widens the range of admissible step-sizes for which the iterates converge to such solutions. These examples provide a rigorous account, aligning with empirical motivation, of how early preconditioners can stabilize and improve dynamics in settings where online SGD fails.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。