arXiv:2410.06186cs.CRcs.LG2024-10ICLR被引 14

提出仅发布最后迭代结果的差分隐私SGD新分析方法,预测更贴近实际隐私泄露。

The Last Iterate Advantage: Empirical Auditing and Principled Heuristic Analysis of Differentially Private SGD

  • 假设模型呈线性结构,仅基于最后迭代估计隐私泄露。
  • 实验验证该方法能准确预测多种训练流程的隐私审计结果。
  • 适合关注真实场景下隐私保护效果的研究者与实践者。

我们提出一种针对仅释放最后迭代结果的噪声裁剪随机梯度下降(DP-SGD)的简单启发式隐私分析。该方法假设模型具有线性结构,可作为训练前对最终隐私泄露的粗略估计。实验表明,该启发式方法能有效预测不同训练流程在隐私审计中的结果,因而具备实用性。同时,我们通过人工反例揭示其局限性:在某些情况下会低估隐私泄露。标准的基于组合的隐私分析假设攻击者可访问所有中间迭代,这在现实中往往不成立,但仍是当前实践中的主流方法。我们的工作展示了理论上限与审计下限之间的巨大差距,并为改进理论分析设定了目标。此外,我们在视觉与语言任务中实证支持该启发式方法,证明现有隐私审计攻击均受其约束。

原文摘要 · Abstract (English)

We propose a simple heuristic privacy analysis of noisy clipped stochastic gradient descent (DP-SGD) in the setting where only the last iterate is released and the intermediate iterates remain hidden. Namely, our heuristic assumes a linear structure for the model. We show experimentally that our heuristic is predictive of the outcome of privacy auditing applied to various training procedures. Thus it can be used prior to training as a rough estimate of the final privacy leakage. We also probe the limitations of our heuristic by providing some artificial counterexamples where it underestimates the privacy leakage. The standard composition-based privacy analysis of DP-SGD effectively assumes that the adversary has access to all intermediate iterates, which is often unrealistic. However, this analysis remains the state of the art in practice. While our heuristic does not replace a rigorous privacy analysis, it illustrates the large gap between the best theoretical upper bounds and the privacy auditing lower bounds and sets a target for further work to improve the theoretical privacy analyses. We also empirically support our heuristic and show existing privacy auditing attacks are bounded by our heuristic analysis in both vision and language tasks.

差分隐私优化算法隐私审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。