提出七项理论突破,让分布式机器学习更高效、更鲁棒、更实用。
Theoretical Foundations of Communication-Efficient, Robust, and Practical Distributed and Federated Optimization

- 通过局部梯度步加速通信,为联邦学习提供理论支撑。
- 在部分设备参与时仍保持通信加速,适应真实场景。
- 首次建立低秩微调的理论框架,助力大模型高效适配。
机器学习与优化协同发展,实际需求推动新理论,理论突破又催生新应用。现代大规模训练依赖经典优化原则,但分布式系统约束要求重新审视理论基础。本文针对联邦学习与分布式优化中的七个关键挑战提出解决方案:首先,提出ProxSkip并证明局部梯度步可加速通信,为广泛使用的启发式方法提供理论依据;其次,设计方差缩减版ProxSkip,消除随机局部更新的邻域误差,平衡通信与计算;第三,在部分客户端参与情况下仍保持通信加速;第四,证明服务器端步长与无放回采样能提升异构环境下的收敛性;第五,发现对梯度差值压缩优于对梯度本身压缩,理论与实践表现更优;第六,证明通过梯度差值裁剪可同时实现拜占庭鲁棒性与部分参与;第七,构建首个基于随机非对称链的低秩适配理论框架,揭示大模型微调的新机制。研究引入新算法框架,建立紧致理论保证,并通过数值实验验证。
原文摘要 · Abstract (English)
Machine learning and optimization have advanced together, with practical demands motivating new theory and theoretical breakthroughs enabling new applications. Modern large-scale training relies on classical optimization principles, but the constraints of distributed systems require these foundations to be reconsidered. This thesis addresses seven challenges at the intersection of theory and practice, focusing on key bottlenecks in federated learning and distributed optimization. First, we introduce ProxSkip and prove that local gradient steps can accelerate communication, providing a theoretical foundation for this widely used heuristic. Second, we develop Variance Reduced ProxSkip, which eliminates the neighborhood error of stochastic local updates while balancing communication and local computation. Third, we show that local steps retain their communication acceleration under partial client participation. Fourth, we prove that server-side stepsizes and sampling without replacement improve convergence in heterogeneous settings. Fifth, for Random Reshuffling, we demonstrate that compressing gradient differences rather than gradients yields better theoretical and practical performance. Sixth, we establish that Byzantine robustness and partial participation can be achieved simultaneously using gradient-difference clipping. Finally, we develop the first theoretical framework for low-rank adaptation based on randomized asymmetric chains, providing new insights into fine-tuning large models. Across these contributions, we introduce novel algorithmic frameworks, establish sharp guarantees under realistic assumptions, and support the theory with numerical experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。