arXiv:2601.10237cs.LGcs.CR2026-01中稿 · ACM CCS 2026被引 3

DP-SGD在强隐私下难以兼顾高精度,噪声要求随训练轮次增长缓慢。

Fundamental Limitations of Favorable Privacy-Utility Guarantees for DP-SGD

  • 在f-差分隐私框架下分析随机梯度下降的隐私-效用权衡机制。
  • 证明了隐私保护越强,所需高斯噪声越大,导致模型精度显著下降。
  • 实验证明该限制在实际训练中已构成性能瓶颈,适合关注隐私计算的研究者阅读。

差分隐私随机梯度下降(DP-SGD)是私有训练的主要方法,但其在最坏情况对抗隐私定义下的根本局限性仍不明确。本文在f-差分隐私框架下分析单个训练周期内经过洗牌采样的DP-SGD,研究了包含M次梯度更新的情形。我们推导出可实现隐私-效用曲线的显式次优上界,并由此得出分离度κ的几何下界。κ表示机制隐私曲线与理想随机猜测线的最大距离,其值越小意味着更强的隐私性。然而我们证明,小κ值会强制要求高斯噪声倍数σ满足σ≥1/√(2ln M),或κ≥1/√8 (1−1/√(4π ln M)),从而严重限制可用效用。尽管当M→∞时噪声阈值趋近于零,但收敛速度极慢;即使在实际训练规模下,所需噪声依然显著。进一步证明该限制对泊松采样也成立,仅差常数因子。实验表明,该边界所隐含的噪声水平会导致实际训练中精度大幅下降,揭示了标准最坏情况对抗假设下DP-SGD的性能瓶颈。

原文摘要 · Abstract (English)

Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant paradigm for private training, but its fundamental limitations under worst-case adversarial privacy definitions remain poorly understood. We analyze DP-SGD in the $f$-differential privacy framework, which characterizes privacy via hypothesis-testing trade-off curves, and study shuffled sampling over a single epoch with $M$ gradient updates. We derive an explicit suboptimal upper bound on the achievable trade-off curve. This result induces a geometric lower bound on the separation $κ$, which is the maximum distance between the mechanism's trade-off curve and the ideal random-guessing line. Because a large separation implies significant adversarial advantage, meaningful privacy requires small $κ$. However, we prove that enforcing a small separation imposes a strict lower bound on the Gaussian noise multiplier $σ$, which directly limits the achievable utility. In particular, under the standard worst-case adversarial model, shuffled DP-SGD must satisfy $$σ\ge \frac{1}{\sqrt{2\ln M}} \quad\text{or}\quad κ\ge\ \frac{1}{\sqrt{8}}\!\left(1-\frac{1}{\sqrt{4π\ln M}}\right),$$ thus cannot simultaneously achieve strong privacy and high utility. Although the noise threshold vanishes asymptotically as $M \to \infty$, the convergence is extremely slow. Even for practically relevant numbers of updates the required noise magnitude remains substantial. We further show that the same limitation extends to Poisson subsampling up to constant factors. Our experiments confirm that the noise levels implied by this bound lead to significant accuracy degradation at realistic training settings, thus showing a bottleneck in DP-SGD under standard worst-case adversarial assumptions.

差分隐私机器学习隐私保护优化瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。