证明了差分隐私随机梯度下降几乎必然收敛,理论更扎实。
Almost Sure Convergence Analysis of Differentially Private Stochastic Gradient Methods
- 在标准衰减步长下,证明了非凸与强凸场景下几乎必然收敛。
- 扩展到带动量的DP-SHB和DP-NAG,保持路径稳定性。
- 为差分隐私优化提供更强理论支撑,适合关注算法稳定性的研究者。
差分隐私随机梯度下降(DP-SGD)已成为实现严格隐私保障的机器学习训练标准算法。尽管广泛应用,其长期行为的理论理解仍有限:现有分析通常仅建立期望收敛或高概率收敛,未涵盖单次轨迹的几乎必然收敛。本文证明,在标准光滑性假设下,只要步长满足常规衰减条件,DP-SGD在非凸与强凸设置中均几乎必然收敛。分析进一步推广至动量变体如随机重球法(DP-SHB)和奈斯特罗夫加速梯度(DP-NAG),通过精心构造能量函数,仍可获得类似保证。结果表明,尽管存在隐私引入的扰动,算法在凸与非凸场景下仍保持路径稳定性,为差分隐私优化提供了更坚实的理论基础。
原文摘要 · Abstract (English)
Differentially private stochastic gradient descent (DP-SGD) has become the standard algorithm for training machine learning models with rigorous privacy guarantees. Despite its widespread use, the theoretical understanding of its long-run behavior remains limited: existing analyses typically establish convergence in expectation or with high probability, but do not address the almost sure convergence of single trajectories. In this work, we prove that DP-SGD converges almost surely under standard smoothness assumptions, both in nonconvex and strongly convex settings, provided the step sizes satisfy some standard decaying conditions. Our analysis extends to momentum variants such as the stochastic heavy ball (DP-SHB) and Nesterov's accelerated gradient (DP-NAG), where we show that careful energy constructions yield similar guarantees. These results provide stronger theoretical foundations for differentially private optimization and suggest that, despite privacy-induced distortions, the algorithm remains pathwise stable in both convex and nonconvex regimes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。