arXiv:2605.10205cs.LG2026-05

提出D-SGD的最优高概率泛化理论,解决分布式学习中的性能差距问题。

Unveiling High-Probability Generalization in Decentralized SGD

  • 用点态一致稳定性分析D-SGD,弱于传统一致稳定性
  • 实现最优泛化率 $\mathcal{O}\left(\frac{1}{\sqrt{mn}}\log (1/δ)\right)$
  • 适用于非凸、强凸等场景,适合研究分布式学习的学者

去中心化随机梯度下降(D-SGD)是大规模分布式学习的有效方法。现有泛化研究多关注期望结果,泛化率受限于 $\mathcal{O}\left(\frac{1}{δ\sqrt{mn}}\right)$,其中 $δ$ 为置信参数,$m$ 为工作节点数,$n$ 为样本量。当 $m=1$ 时,D-SGD退化为传统 SGD,其最优高概率泛化界为 $\mathcal{O}\left(\frac{1}{\sqrt{n}}\log (1/δ)\right)$,二者存在差距。本文旨在填补该空白,构建 D-SGD 的高概率学习理论,目标达到最优 $\mathcal{O}\left(\frac{1}{\sqrt{mn}}\log (1/δ)\right)$ 率。通过引入分布式学习中的点态一致稳定性(弱于一致稳定性),在凸、强凸及非凸设置下重构边界,并针对仅存在局部极小值的非凸情形,给出基于梯度的高概率结果,推导优化误差与过拟合风险界。最后,考虑通信开销,在时变框架下分析本地模型的泛化界。

原文摘要 · Abstract (English)

Decentralized stochastic gradient descent (D-SGD) is an efficient method for large-scale distributed learning. Existing generalization studies mainly address expected results, achieving rates limited to $\mathcal{O}\left(\frac{1}{δ\sqrt{mn}}\right)$, where $δ$ is the confidence parameter, $m$ the number of workers, and $n$ the sample size. When $m=1$, D-SGD reduces to traditional SGD, whose optimal high-probability generalization bound is $\mathcal{O}\left(\frac{1}{\sqrt{n}}\log (1/δ)\right)$. This discrepancy reveals a gap between high-probability guarantees for SGD and those for D-SGD. To close this, we develop a high-probability learning theory for D-SGD, aiming for the optimal $\mathcal{O}\left(\frac{1}{\sqrt{mn}}\log (1/δ)\right)$ rate. We refine bounds for D-SGD using pointwise uniform stability in distributed learning-a weaker notion than uniform stability-and analyze them across convex, strongly convex, and non-convex settings. We also provide high-probability results for gradient-based measures in non-convex cases where only local minima exist, and derive optimization error and excess risk bounds. Finally, accounting for communication overhead, we analyze generalization bounds for local models within time-varying frameworks.

分布式学习泛化理论D-SGD

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。