发现排序采样比泊松采样隐私保护弱,质疑主流方法的隐私评估
Scalable DP-SGD: Shuffling vs. Poisson Subsampling
- 对比分析多轮排序与泊松采样在隐私上的差异
- 证明排序采样实际隐私预算远差于报告值,差距显著
- 提出可扩展的泊松采样实现方案,支持高效训练
我们为多轮自适应批量线性查询(ABLQ)机制在排序批量采样下的隐私保证提供了新的下界,发现其与泊松采样相比存在显著差距;此前分析仅限单轮。由于差分隐私随机梯度下降(DP-SGD)的隐私分析基于ABLQ机制,这使得以排序采样实现但报告为泊松采样的常见做法受到严重质疑。为评估该差距对模型性能的影响,我们提出一种基于大规模并行计算的可扩展泊松采样实现方法,并高效训练模型。通过新下界对比了基于泊松采样的模型与使用排序采样时乐观估计的模型性能。
原文摘要 · Abstract (English)
We provide new lower bounds on the privacy guarantee of the multi-epoch Adaptive Batch Linear Queries (ABLQ) mechanism with shuffled batch sampling, demonstrating substantial gaps when compared to Poisson subsampling; prior analysis was limited to a single epoch. Since the privacy analysis of Differentially Private Stochastic Gradient Descent (DP-SGD) is obtained by analyzing the ABLQ mechanism, this brings into serious question the common practice of implementing shuffling-based DP-SGD, but reporting privacy parameters as if Poisson subsampling was used. To understand the impact of this gap on the utility of trained machine learning models, we introduce a practical approach to implement Poisson subsampling at scale using massively parallel computation, and efficiently train models with the same. We compare the utility of models trained with Poisson-subsampling-based DP-SGD, and the optimistic estimates of utility when using shuffling, via our new lower bounds on the privacy guarantee of ABLQ with shuffling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。