arXiv:2412.19916cs.LGcs.CR2024-12被引 8

首次理论分析自适应梯度裁剪的收敛性,解决隐私优化中的参数敏感问题。

On the Convergence of DP-SGD with Adaptive Clipping

  • 提出基于分位数的自适应裁剪方法,避免固定阈值依赖
  • 证明分位数选择与学习率需协同设计才能保证收敛
  • 为实际应用提供可操作的参数配置指导,适合隐私机器学习研究者

带有梯度裁剪的随机梯度下降(SGD)是实现差分隐私优化的有效技术。尽管已有大量研究关注固定阈值裁剪,但隐私训练对阈值选择极为敏感,调参成本高甚至不可行。这促使了自适应方法的发展,如分位数裁剪(QC-SGD),其在实践中表现良好,但缺乏坚实的理论支撑。本文首次对QC-SGD进行了全面的收敛性分析,发现其存在与固定阈值裁剪类似的偏差问题,但通过精心设计的分位数与步长调度可有效缓解。分析揭示了分位数选择、步长与收敛行为之间的关键关系,提供了实用的参数选取指南。进一步将结果扩展至差分隐私优化,首次建立了DP-QC-SGD的理论保证。研究为广泛使用的自适应裁剪启发式方法提供了理论基础,并指明了未来研究方向。

原文摘要 · Abstract (English)

Stochastic Gradient Descent (SGD) with gradient clipping is a powerful technique for enabling differentially private optimization. Although prior works extensively investigated clipping with a constant threshold, private training remains highly sensitive to threshold selection, which can be expensive or even infeasible to tune. This sensitivity motivates the development of adaptive approaches, such as quantile clipping, which have demonstrated empirical success but lack a solid theoretical understanding. This paper provides the first comprehensive convergence analysis of SGD with quantile clipping (QC-SGD). We demonstrate that QC-SGD suffers from a bias problem similar to constant-threshold clipped SGD but show how this can be mitigated through a carefully designed quantile and step size schedule. Our analysis reveals crucial relationships between quantile selection, step size, and convergence behavior, providing practical guidelines for parameter selection. We extend these results to differentially private optimization, establishing the first theoretical guarantees for DP-QC-SGD. Our findings provide theoretical foundations for widely used adaptive clipping heuristic and highlight open avenues for future research.

差分隐私自适应裁剪梯度下降收敛性分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。