arXiv:2506.00181cs.LGstat.ML2025-06中稿 · ICML被引 4

揭示分布式优化中噪声、压缩与自适应的相互作用机制

On the Interaction of Batch Noise, Adaptivity, and Compression, under $(L_0,L_1)$-Smoothness: An SDE Approach

  • 用随机微分方程建模分布式优化中的梯度噪声与压缩影响
  • 发现自适应更新需根据噪声结构和压缩率动态调整以保证稳定
  • 适用于研究分布式训练稳定性或设计鲁棒优化算法的研究者

分布式随机优化同时涉及(一)随机梯度噪声,(二)通信压缩,(三)自适应/归一化更新。尽管各因素已分别研究,其在真实假设下的联合效应仍不清晰。本文针对最近提出的$(L_0, L_1)$-smoothness条件,构建了分布式压缩SGD(DCSGD)及其符号变体(DSignSGD)的统一理论框架。从概念上,指出文献中的一阶与二阶修正方程无法准确刻画离散时间步长/稳定性限制,尤其在$(L_0,L_1)$-smoothness下。技术上,通过将曲率相关项精确嵌入漂移项,提出新的首阶随机微分方程,从而捕捉学习率限制、梯度噪声、压缩率与损失曲面几何之间的精细关系。关键在于,模型在一般梯度噪声假设下成立,包括重尾和仿射方差情形,超越经典有界方差设定。结果表明,DCSGD的更新归一化是稳定性的自然条件,其程度由噪声结构、曲面正则性与压缩率共同决定;而DSignSGD在重尾噪声下仍可用标准学习率收敛。这些发现提供新理论洞见与实用指导。

原文摘要 · Abstract (English)

Distributed stochastic optimization intertwines (i) stochastic gradient noise, (ii) communication compression, and (iii) adaptive/normalized updates. While each factor has been studied in isolation, their joint effect under realistic assumptions remains poorly understood. In this work, we develop a unified theoretical framework for Distributed Compressed SGD (DCSGD) and its sign variant Distributed SignSGD (DSignSGD) under the recently introduced $(L_0, L_1)$-smoothness condition. From a conceptual perspective, we show that the first- and second-order modified equations from the literature do not accurately model the discrete-time step-size/stability restrictions, especially under $(L_0,L_1)$-smoothness. From a technical perspective, we propose new first-order SDEs by carefully incorporating curvature-dependent terms into their drift: This helps capture the fine-grained relationship between learning rate restrictions, gradient noise, compression, and the geometry of the loss landscape. Importantly, we do so under general gradient noise assumptions, including heavy-tailed and affine-variance regimes, which extend beyond the classical bounded-variance setting. Our results suggest that normalizing the updates of DCSGD emerges as a natural condition for stability, with the degree of normalization precisely determined by the gradient noise structure, the landscape's regularity, and the compression rate. In contrast, DSignSGD converges even under heavy-tailed noise with standard learning rate schedules. Together, these findings offer both new theoretical insights and perspectives, and practical guidance.

分布式优化随机微分方程压缩自适应学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。