arXiv:2603.03226cs.LGcs.CR2026-03中稿 · ICLR被引 2

高隐私场景下自适应优化器更优,因噪声与学习率适配性更强。

Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective

  • 从随机微分方程视角分析私有优化器的噪声与自适应关系。
  • 在高隐私设置下,DP-SignSGD收敛速度随ε线性提升,优于DP-SGD。
  • 自适应方法可跨隐私级别复用超参数,实操性显著更强。

随着隐私监管趋严,差分隐私(DP)正成为大规模训练的核心。本文通过随机微分方程(SDE)视角重新审视DP噪声与优化自适应性的交互关系,首次对私有优化器进行基于SDE的分析。聚焦于采用逐样本裁剪的DP-SGD与DP-SignSGD,发现固定超参数下:DP-SGD的隐私-效用权衡为$/mathcal{O}(1/\varepsilon^2)$,收敛速度与ε无关;而DP-SignSGD收敛速度与ε线性相关,权衡为$/mathcal{O}(1/\varepsilon)$,在高隐私或大批次噪声场景中占优。当使用最优学习率时,两者理论渐近性能相当,但DP-SGD的最优学习率与ε线性相关,而DP-SignSGD几乎与ε无关。这使得自适应方法更实用,因其超参数可在不同隐私水平间复用,无需大量调参。实验结果在训练与测试指标上验证了理论,并将结论扩展至DP-Adam。

原文摘要 · Abstract (English)

Differential Privacy (DP) is becoming central to large-scale training as privacy regulations tighten. We revisit how DP noise interacts with adaptivity in optimization through the lens of stochastic differential equations, providing the first SDE-based analysis of private optimizers. Focusing on DP-SGD and DP-SignSGD under per-example clipping, we show a sharp contrast under fixed hyperparameters: DP-SGD converges at a Privacy-Utility Trade-Off of $\mathcal{O}(1/\varepsilon^2)$ with speed independent of $\varepsilon$, while DP-SignSGD converges at a speed linear in $\varepsilon$ with an $\mathcal{O}(1/\varepsilon)$ trade-off, dominating in high-privacy or large batch noise regimes. By contrast, under optimal learning rates, both methods achieve comparable theoretical asymptotic performance; however, the optimal learning rate of DP-SGD scales linearly with $\varepsilon$, while that of DP-SignSGD is essentially $\varepsilon$-independent. This makes adaptive methods far more practical, as their hyperparameters transfer across privacy levels with little or no re-tuning. Empirical results confirm our theory across training and test metrics, and empirically extend from DP-SignSGD to DP-Adam.

差分隐私自适应优化模型训练随机微分方程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。