提出加权聚合方法,提升异步分布式训练的抗错能力。
Weight for Robustness: A Comprehensive Approach towards Optimal Fault-Tolerant Asynchronous ML
- 设计加权鲁棒聚合框架,缓解延迟更新影响。
- 首次实现异步拜占庭环境下的最优收敛速度。
- 适合研究分布式机器学习容错与高效训练的人群。
针对异步分布式机器学习系统中拜占庭鲁棒性训练的挑战,本文旨在提升大规模并行化和异构计算资源下的训练效率。异步系统中,各工作节点独立运行且更新周期不固定,对拜占庭故障(恶意或错误行为)的抵御能力显著下降,固有延迟不仅引入额外偏差,还掩盖了故障造成的干扰。为此,本文将拜占庭框架适配至异步动态,提出新型加权鲁棒聚合机制,使现有鲁棒聚合器与近期元聚合器可扩展为加权版本,有效缓解延迟更新的影响。进一步结合最新方差减少技术,首次在异步拜占庭环境中实现最优收敛率。通过理论与实验验证,该方法显著增强系统容错性并优化性能。
原文摘要 · Abstract (English)
We address the challenges of Byzantine-robust training in asynchronous distributed machine learning systems, aiming to enhance efficiency amid massive parallelization and heterogeneous computing resources. Asynchronous systems, marked by independently operating workers and intermittent updates, uniquely struggle with maintaining integrity against Byzantine failures, which encompass malicious or erroneous actions that disrupt learning. The inherent delays in such settings not only introduce additional bias to the system but also obscure the disruptions caused by Byzantine faults. To tackle these issues, we adapt the Byzantine framework to asynchronous dynamics by introducing a novel weighted robust aggregation framework. This allows for the extension of robust aggregators and a recent meta-aggregator to their weighted versions, mitigating the effects of delayed updates. By further incorporating a recent variance-reduction technique, we achieve an optimal convergence rate for the first time in an asynchronous Byzantine environment. Our methodology is rigorously validated through empirical and theoretical analysis, demonstrating its effectiveness in enhancing fault tolerance and optimizing performance in asynchronous ML systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。