同步SGD在异构计算中近乎最优,挑战了异步方法的必要性。
Do We Need Asynchronous SGD? On the Near-Optimality of Synchronous Methods
- 分析随机计算时间与部分工作节点参与下的同步SGD
- 证明其时间复杂度在多数场景下接近理论最优,仅差对数因子
- 适合异构计算环境,尤其适用于资源不均的分布式训练
现代分布式优化大多依赖传统的同步方法,尽管异步优化近年取得显著进展。本文重新审视同步SGD及其鲁棒变体m-同步SGD,理论证明它们在多种异构计算场景下近乎最优,这一结果出乎意料。在随机计算时间与对抗性部分参与条件下,我们证明同步方法的时间复杂度在多数实际情形下为最优,仅相差对数因子。虽然同步方法并非万能,某些任务仍需异步方法,但其已足够应对大量现代异构计算场景。
原文摘要 · Abstract (English)
Modern distributed optimization methods mostly rely on traditional synchronous approaches, despite substantial recent progress in asynchronous optimization. We revisit Synchronous SGD and its robust variant, called $m$-Synchronous SGD, and theoretically show that they are nearly optimal in many heterogeneous computation scenarios, which is somewhat unexpected. We analyze the synchronous methods under random computation times and adversarial partial participation of workers, and prove that their time complexities are optimal in many practical regimes, up to logarithmic factors. While synchronous methods are not universal solutions and there exist tasks where asynchronous methods may be necessary, we show that they are sufficient for many modern heterogeneous computation scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。