突破传统平滑性假设,实现有保证的分布式优化收敛。
Provably Convergent Decentralized Optimization over Directed Graphs under Generalized Smoothness
- 引入广义光滑性框架,适应梯度快速变化场景。
- 在无界梯度差异下仍保持收敛,优于现有方法。
- 适用于真实异构数据,适合大规模分布式学习系统。
去中心化优化已成为大规模学习系统的核心工具;然而,现有方法多依赖经典Lipschitz光滑性假设,该假设在梯度快速变化的问题中常被违反。为克服此局限,本文研究在广义$(L_0, L_1)$-光滑性框架下的去中心化优化,其中海森矩阵范数允许随梯度范数线性增长,从而容纳超出经典Lipschitz光滑性的快速变化梯度。通过结合梯度追踪与梯度裁剪,并精心设计裁剪阈值,确保在有向通信图下实现精确收敛。与现有在广义光滑性下需有界梯度差异假设的结果不同,本方法在梯度差异无界时依然有效,使框架更适用于现实异构数据环境。我们在LIBSVM和CIFAR-10等标准数据集上,采用正则化逻辑回归和卷积神经网络进行数值实验,验证了所提方法在稳定性与收敛速度上均优于现有方法。
原文摘要 · Abstract (English)
Decentralized optimization has become a fundamental tool for large-scale learning systems; however, most existing methods rely on the classical Lipschitz smoothness assumption, which is often violated in problems with rapidly varying gradients. Motivated by this limitation, we study decentralized optimization under the generalized $(L_0, L_1)$-smoothness framework, in which the Hessian norm is allowed to grow linearly with the gradient norm, thereby accommodating rapidly varying gradients beyond classical Lipschitz smoothness. We integrate gradient-tracking techniques with gradient clipping and carefully design the clipping threshold to ensure accurate convergence over directed communication graphs under generalized smoothness. In contrast to existing distributed optimization results under generalized smoothness that require a bounded gradient dissimilarity assumption, our results remain valid even when the gradient dissimilarity is unbounded, making the proposed framework more applicable to realistic heterogeneous data environments. We validate our approach via numerical experiments on standard benchmark datasets, including LIBSVM and CIFAR-10, using regularized logistic regression and convolutional neural networks, demonstrating superior stability and faster convergence over existing methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。