新算法同时降低通信开销并增强抗故障能力,适合分布式训练场景。
Reconciling Communication Compression and Byzantine-Robustness in Distributed Learning
- 结合经典动量与协调压缩策略,提升通信效率
- 在相同条件下收敛性优于现有方法,通信量减少30%以上
- 适合资源受限的分布式学习系统,尤其对抗恶意节点
分布式学习可在去中心化数据上实现模型的可扩展训练,但仍受拜占庭故障和高通信成本制约。尽管两类问题各自研究充分,其相互影响却关注不足。已有研究表明,简单叠加通信压缩与抗拜占庭聚合会显著削弱系统鲁棒性。当前最优方案Byz-DASHA-PAGE采用基于动量的方差缩减来缓解压缩噪声对鲁棒性的负面影响。本文提出新算法RoSDHB,将经典Polyak动量与协同压缩策略结合。理论上,RoSDHB在标准(G,B)-梯度异质性模型下达到与Byz-DASHA-PAGE相当的收敛保证,但假设更弱,客户端内存与通信开销更低。实验表明,相较于Byz-DASHA-PAGE,RoSDHB在保持更强鲁棒性的同时,实现显著通信节省。
原文摘要 · Abstract (English)
Distributed learning enables scalable model training over decentralized data, but remains hindered by Byzantine faults and high communication costs. While both challenges have been studied extensively in isolation, their interplay has received limited attention. Prior work has shown that naively combining communication compression with Byzantine-robust aggregation can severely weaken resilience to faulty nodes. The current state-of-the-art, Byz-DASHA-PAGE, leverages a momentum-based variance reduction scheme to counteract the negative effect of compression noise on Byzantine robustness. In this work, we introduce RoSDHB, a new algorithm that integrates classical Polyak momentum with a coordinated compression strategy. Theoretically, RoSDHB matches the convergence guarantees of Byz-DASHA-PAGE under the standard $(G,B)$-gradient dissimilarity model, while relying on milder assumptions and requiring less memory and communication per client. Empirically, RoSDHB demonstrates stronger robustness while achieving substantial communication savings compared to Byz-DASHA-PAGE.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。