解析signSGD在高维下的风险曲线,量化其降噪与预处理效应。
Exact Risk Curves of signSGD in High-Dimensions: Quantifying Preconditioning and Noise-Compression Effects
- 通过高维极限推导出描述风险的随机微分方程和常微分方程
- 定量揭示了有效学习率、噪声压缩、对角预处理和梯度噪声重塑四类效应
- 适用于理解Adam等自适应优化器,为理论分析提供新框架
近年来,signSGD因其作为实用优化器及研究自适应优化器(如Adam)的简化模型而受到关注。尽管普遍认为signSGD具有预处理和噪声重塑作用,但在理论上可解的设置下对其效应进行量化仍具挑战。本文在高维极限下分析signSGD,推导出描述风险的极限随机微分方程(SDE)和常微分方程(ODE)。基于此框架,我们量化了signSGD的四个效应:有效学习率、噪声压缩、对角预处理和梯度噪声重塑。分析结果与实验观察一致,并进一步揭示这些效应如何依赖于数据和噪声分布。最后,我们提出一个关于如何将这些结果扩展至Adam的猜想。
原文摘要 · Abstract (English)
In recent years, signSGD has garnered interest as both a practical optimizer as well as a simple model to understand adaptive optimizers like Adam. Though there is a general consensus that signSGD acts to precondition optimization and reshapes noise, quantitatively understanding these effects in theoretically solvable settings remains difficult. We present an analysis of signSGD in a high dimensional limit, and derive a limiting SDE and ODE to describe the risk. Using this framework we quantify four effects of signSGD: effective learning rate, noise compression, diagonal preconditioning, and gradient noise reshaping. Our analysis is consistent with experimental observations but moves beyond that by quantifying the dependence of these effects on the data and noise distributions. We conclude with a conjecture on how these results might be extended to Adam.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。