提出通信高效的联邦学习优化算法,显著降低模型训练通信开销。
Communication-Efficient Gluon in Federated Learning
- 基于层间平滑性设计压缩优化框架,结合方差缩减技术提升精度。
- 在多种条件下实现比现有方法更低的通信成本,且收敛更快。
- 适合大规模分布式训练场景,尤其适用于资源受限的设备部署。
近期研究显示,基于非欧几里得范数球上线性最小化预言机(LMO)的Muon型优化器在大语言模型训练中表现优于Adam类方法。由于大规模神经网络需在海量机器上训练,通信开销成为瓶颈。为此,本文研究Gluon——Muon在更通用的逐层$(L^0, L^1)$-光滑设置下的扩展,采用无偏与收缩压缩器。为降低压缩误差,引入SARAH中的方差缩减技术。在特定条件下,实现了更优的收敛速率与通信效率。作为副产品,获得一种比Gluon收敛更快的新方差缩减算法。进一步结合动量方差缩减(MVR),当$L_i^1 \neq 0$时,在更弱条件下仍保持相当的通信成本。多个数值实验验证了所提压缩算法在通信开销上的优越性能。
原文摘要 · Abstract (English)
Recent developments have shown that Muon-type optimizers based on linear minimization oracles (LMOs) over non-Euclidean norm balls have the potential to get superior practical performance than Adam-type methods in the training of large language models. Since large-scale neural networks are trained across massive machines, communication cost becomes the bottleneck. To address this bottleneck, we investigate Gluon, which is an extension of Muon under the more general layer-wise $(L^0, L^1)$-smooth setting, with both unbiased and contraction compressors. In order to reduce the compression error, we employ the variance reduced technique in SARAH in our compressed methods. The convergence rates and improved communication cost are achieved under certain conditions. As a byproduct, a new variance reduced algorithm with faster convergence rate than Gluon is obtained. We also incorporate momentum variance reduction (MVR) to these compressed algorithms and comparable communication cost is derived under weaker conditions when $L_i^1 \neq 0$. Finally, several numerical experiments are conducted to verify the superior performance of our compressed algorithms in terms of communication cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。