最后一轮全局合并能显著提升去中心化学习性能,尤其在数据异构时。
On the Surprising Effectiveness of a Single Global Merging in Decentralized Learning
- 训练后期集中进行一次全局模型合并,取代频繁通信。
- 在高数据异构下,单次全局合并使性能接近并行训练收敛速度。
- 适合关注通信效率与去中心化模型融合的研究者。
去中心化学习提供了一种可扩展的替代方案,但其性能常受限于有限的设备间通信。本文研究通信调度策略,发现将通信资源集中在训练后期能显著提升全局测试性能。令人意外的是,在高数据异构条件下,仅在最后一步执行全连接通信(即单次全局合并)即可大幅改善去中心化学习表现。理论分析首次证明,去中心化SGD的全局合并模型可达到与并行SGD相同的收敛速率。技术上,我们重新解释了局部模型间的差异——此前被视为有害噪声,实则为实现该速率的关键构造成分。本工作表明,去中心化学习在高异构性与低通信条件下仍具泛化能力,并为模型合并研究开辟新方向。
原文摘要 · Abstract (English)
Decentralized learning provides a scalable alternative to parameter-server-based training, yet its performance is often hindered by limited peer-to-peer communication. In this paper, we study how communication should be scheduled over time, including determining when and how frequently devices synchronize. Counterintuitive empirical results show that concentrating communication budgets in the later stages of decentralized training remarkably improves global test performance. Surprisingly, we uncover that fully connected communication at the final step, implemented by a single global merging, can significantly improve the performance of decentralized learning under high data heterogeneity. Our theoretical contributions, which explain these phenomena, are the first to establish that the globally merged model of decentralized SGD can match the convergence rate of parallel SGD. Technically, we reinterpret part of the discrepancy among local models, which were previously considered as detrimental noise, as constructive components essential for matching this rate. This work provides evidence that decentralized learning is able to generalize under high data heterogeneity and limited communication, while offering broad new avenues for model merging research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。