arXiv:2412.01756cs.CRcs.LG2024-12被引 1

通过对抗样本生成,更精准评估私有模型的隐私泄露风险。

Adversarial Sample-Based Approach for Tighter Privacy Auditing in Final Model-Only Scenarios

  • 用损失函数构造最坏情况输入,提升审计精度。
  • 在MNIST上实测隐私泄露下限达4.914,优于基线4.385。
  • 适合关注模型隐私安全的研究者和开发者使用。

在仅能访问最终模型的场景下,对差分隐私随机梯度下降(DP-SGD)进行隐私审计面临挑战,通常得到的实证下界远低于理论保证。本文提出一种新审计方法,通过基于损失的输入空间审计生成对抗性最坏样本,在不引入额外假设的情况下实现更紧的实证下界。该方法超越传统基于金丝雀的启发式策略,适用于最终模型仅可访问场景。具体而言,在理论隐私预算 ε = 10.0 的条件下,本方法在MNIST数据集上实现了4.914的实证下界,优于基线的4.385。本工作为差分隐私机器学习中的可靠、准确隐私审计提供了实用框架。

原文摘要 · Abstract (English)

Auditing Differentially Private Stochastic Gradient Descent (DP-SGD) in the final model setting is challenging and often results in empirical lower bounds that are significantly looser than theoretical privacy guarantees. We introduce a novel auditing method that achieves tighter empirical lower bounds without additional assumptions by crafting worst-case adversarial samples through loss-based input-space auditing. Our approach surpasses traditional canary-based heuristics and is effective in final model-only scenarios. Specifically, with a theoretical privacy budget of $\varepsilon = 10.0$, our method achieves empirical lower bounds of $4.914$, compared to the baseline of $4.385$ for MNIST. Our work offers a practical framework for reliable and accurate privacy auditing in differentially private machine learning.

隐私审计差分隐私对抗样本模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。