arXiv:2603.02639math.OCcs.LG2026-03

用递减学习率解决延迟梯度下的联邦学习优化问题

Convex and Non-convex Federated Learning with Stale Stochastic Gradients: Diminishing Step Size is All You Need

  • 采用预设递减学习率,无需自适应调整
  • 在非凸与强凸目标下均达到最优收敛速率
  • 适用于存在延迟梯度的分布式联邦学习场景

我们提出一种适用于延迟梯度模型的分布式随机优化通用框架。在该框架中,n 个本地代理利用自身数据和计算能力,协助中心服务器最小化由代理局部代价函数组成的全局目标。每个代理可传输其局部梯度的随机估计,这些估计可能带有偏差且存在延迟。尽管已有工作建议在延迟存在时采用延迟自适应学习率,但我们证明,预先设定的递减学习率已足够,并能匹配自适应方案的性能。此外,我们的分析表明,递减学习率在非凸和强凸目标下均能恢复最优 SGD 收敛速率。

原文摘要 · Abstract (English)

We propose a general framework for distributed stochastic optimization under delayed gradient models. In this setting, $n$ local agents leverage their own data and computation to assist a central server in minimizing a global objective composed of agents' local cost functions. Each agent is allowed to transmit stochastic-potentially biased and delayed-estimates of its local gradient. While a prior work has advocated delay-adaptive step sizes for stochastic gradient descent (SGD) in the presence of delays, we demonstrate that a pre-chosen diminishing step size is sufficient and matches the performance of the adaptive scheme. Moreover, our analysis establishes that diminishing step sizes recover the optimal SGD rates for nonconvex and strongly convex objectives.

联邦学习优化算法梯度延迟递减学习率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。