改进隐私保护训练方法,提升模型准确率同时保证严格隐私。
Revisiting Privacy Amplification by Subsampling in Selective Release DPSGD

- 基于梯度裁剪设计新算法,重新分析选择性释放的隐私增益。
- 在多个数据集上实现高精度,如MNIST上达到99.1%准确率。
- 适合需要强隐私保障又追求高性能的机器学习应用。
机器学习依赖敏感数据,需采用差分隐私随机梯度下降(DPSGD)等隐私保护技术。但DPSGD因梯度裁剪和噪声注入导致性能下降、收敛缓慢。已有研究尝试优化,其中差分隐私选择性更新与释放(DPSUR)算法显著提升了模型性能。然而,其隐私分析忽略了选择性释放带来的采样概率变化,削弱了隐私保证的严谨性。为此,本文重新评估选择性释放机制的隐私分析,提出新算法:基于裁剪梯度的差分隐私选择性释放(DPSR-CG)。通过严格的全新隐私分析及在MNIST、CIFAR-10、IMDB和FMNIST等多个数据集上的实验验证,表明DPSR-CG在保持严格隐私保障的同时,实现了卓越的模型表现。
原文摘要 · Abstract (English)
Machine learning's reliance on sensitive data necessitates privacy-preserving techniques like Differentially Private Stochastic Gradient Descent (DPSGD). However, DPSGD suffers from substantial utility degradation and slow convergence due to gradient clipping and noise injection. Prior works have attempted to improve DPSGD from various perspectives; notably, the Differentially Private Selective Update and Release (DPSUR) algorithm has achieved remarkable model utility. However, the privacy accounting in DPSUR overlooks the variation in sampling probability introduced by the selective release mechanism, which compromises the rigor of its privacy guarantees. To address these limitations, we re-evaluate the privacy analysis of the selective release mechanism and propose a novel algorithm: Differentially Private Selective Release based on Clipped Gradients (DPSR-CG). Through a rigorous, newly derived privacy analysis and extensive experiments on multiple datasets (MNIST, CIFAR-10, IMDB, and FMNIST), we demonstrate that our DPSR-CG mechanism maintains strict privacy guarantees while achieving exceptional model performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。