为隐私保护的梯度下降设计了有效的统计推断方法。
Statistical Inference for Differentially Private Stochastic Gradient Descent
- 提出随机采样下SGD的渐近性质,拓展至DP-SGD。
- DP-SGD输出方差分解为统计、采样与隐私三部分。
- 两种置信区间方法在保持隐私前提下覆盖率达预期。
机器学习中的隐私保护,尤其是通过差分隐私随机梯度下降(DP-SGD)在敏感数据分析中至关重要。然而,现有针对SGD的统计推断方法主要聚焦于循环采样,而DP-SGD需采用随机采样。本文首次填补这一空白,建立了随机采样规则下SGD的渐近性质,并将其推广至DP-SGD。对于DP-SGD的输出,我们证明其渐近方差可分解为统计、采样和隐私诱导三部分。本文提出两种构建有效置信区间的方案:插件法与随机缩放法。通过广泛的数值分析表明,所提置信区间在保持隐私的同时实现了名义覆盖率。
原文摘要 · Abstract (English)
Privacy preservation in machine learning, particularly through Differentially Private Stochastic Gradient Descent (DP-SGD), is critical for sensitive data analysis. However, existing statistical inference methods for SGD predominantly focus on cyclic subsampling, while DP-SGD requires randomized subsampling. This paper first bridges this gap by establishing the asymptotic properties of SGD under the randomized rule and extending these results to DP-SGD. For the output of DP-SGD, we show that the asymptotic variance decomposes into statistical, sampling, and privacy-induced components. Two methods are proposed for constructing valid confidence intervals: the plug-in method and the random scaling method. We also perform extensive numerical analysis, which shows that the proposed confidence intervals achieve nominal coverage rates while maintaining privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。