改进了差分隐私SGD的隐私与性能权衡分析,无需凸性假设。
An Improved Privacy and Utility Analysis of Differentially Private SGD with Bounded Domain and Smooth Losses
- 基于平滑损失和有界域,追踪多轮隐私损耗变化
- 证明无凸性时隐私损耗仍可收敛,且更小域直径提升隐私与性能
- 揭示了隐私-效用权衡的最优阶数,适合关注数据安全的模型训练者
差分隐私随机梯度下降(DPSGD)广泛用于保护机器学习训练中的敏感数据,但其隐私保障常以模型性能大幅下降为代价,原因在于缺乏紧致的隐私损失理论边界。尽管近期研究已实现更精确的隐私保证,但仍依赖实际中不适用的假设,如凸性及复杂参数要求,且很少深入探讨隐私机制对模型效用的影响。本文针对一般L-平滑、非凸损失函数下的DPSGD,提供了严格的隐私表征,揭示了在有界域情况下隐私损失随迭代收敛的规律。具体而言,通过利用噪声平滑递减特性,跟踪多轮隐私损耗,并在不同场景下建立全面的收敛性分析。结果表明:(i) 在有界域下,即使无凸性假设,隐私损失仍可收敛;(ii) 在特定条件下,更小的有界直径能同时提升隐私与模型效用;(iii) 对于带梯度裁剪的DPSGD(DPSGD-GC)及其有界域版本(DPSGD-DC),在强凸总体风险函数下,分别实现了可达的隐私-效用权衡大O阶数。通过真实场景下的成员推断攻击(MIA)实验验证了理论洞察。
原文摘要 · Abstract (English)
Differentially Private Stochastic Gradient Descent (DPSGD) is widely used to protect sensitive data during the training of machine learning models, but its privacy guarantee often comes at a large cost of model performance due to the lack of tight theoretical bounds quantifying privacy loss. While recent efforts have achieved more accurate privacy guarantees, they still impose some assumptions prohibited from practical applications, such as convexity and complex parameter requirements, and rarely investigate in-depth the impact of privacy mechanisms on the model's utility. In this paper, we provide a rigorous privacy characterization for DPSGD with general L-smooth and non-convex loss functions, revealing converged privacy loss with iteration in bounded-domain cases. Specifically, we track the privacy loss over multiple iterations, leveraging the noisy smooth-reduction property, and further establish comprehensive convergence analysis in different scenarios. In particular, we show that for DPSGD with a bounded domain, (i) the privacy loss can still converge without the convexity assumption, (ii) a smaller bounded diameter can improve both privacy and utility simultaneously under certain conditions, and (iii) the attainable big-O order of the privacy utility trade-off for DPSGD with gradient clipping (DPSGD-GC) and for DPSGD-GC with bounded domain (DPSGD-DC) and mu-strongly convex population risk function, respectively. Experiments via membership inference attack (MIA) in a practical setting validate insights gained from the theoretical results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。