arXiv:2507.06969cs.LGcs.AI2025-07NeurIPS被引 13

统一了差分隐私中三类风险的评估方法,让防护更精准高效。

Unifying Re-Identification, Attribute Inference, and Data Reconstruction Risks in Differential Privacy

  • 基于假设检验重构差分隐私,建立三类攻击风险的统一分析框架。
  • 实测比现有方法更紧致,相同风险下可降低20%噪声,文本分类准确率提升18个百分点。
  • 适合需要精确量化隐私风险的开发者与研究人员使用。

差分隐私(DP)机制难以解释和校准,因为现有方法在将标准隐私参数映射到具体隐私风险(再识别、属性推断、数据重建)时,既过于悲观又不一致。本文采用差分隐私的假设检验解释(f-DP),发现三类风险的攻击成功率边界可统一为相同形式。该统一边界具备一致性,适用于多种攻击场景,并且可调节,使从业者能针对任意(包括最坏情况)基线风险水平评估风险。实验表明,本方法所得边界比传统ε-DP、Rényi DP及集中式DP更紧致。使用该边界校准噪声,可在相同风险水平下减少20%噪声,例如在文本分类任务中使准确率从52%提升至70%。总体而言,这一统一视角为评估和校准差分隐私对特定再识别、属性推断或数据重建风险的保护程度提供了原则性框架。

原文摘要 · Abstract (English)

Differentially private (DP) mechanisms are difficult to interpret and calibrate because existing methods for mapping standard privacy parameters to concrete privacy risks -- re-identification, attribute inference, and data reconstruction -- are both overly pessimistic and inconsistent. In this work, we use the hypothesis-testing interpretation of DP ($f$-DP), and determine that bounds on attack success can take the same unified form across re-identification, attribute inference, and data reconstruction risks. Our unified bounds are (1) consistent across a multitude of attack settings, and (2) tunable, enabling practitioners to evaluate risk with respect to arbitrary, including worst-case, levels of baseline risk. Empirically, our results are tighter than prior methods using $\varepsilon$-DP, Rényi DP, and concentrated DP. As a result, calibrating noise using our bounds can reduce the required noise by 20% at the same risk level, which yields, e.g., an accuracy increase from 52% to 70% in a text classification task. Overall, this unifying perspective provides a principled framework for interpreting and calibrating the degree of protection in DP against specific levels of re-identification, attribute inference, or data reconstruction risk.

差分隐私风险评估统一框架隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。