让少数特殊个体更安全:通过降低其数据影响来强化差分隐私保护
Risk-Equalized Differentially Private Synthetic Data: Protecting Outliers by Controlling Record-Level Influence
- 先评估每条数据的异常程度,再按风险高低加权调整其对生成模型的影响
- 实验显示对高异常值的成员推断攻击成功率显著下降,最高降幅达47%
- 适合关注隐私保护公平性的医疗、金融等敏感数据合成场景
当发布合成数据时,某些个体(如罕见病患者或异常交易)比其他人更难保护。尽管差分隐私提供最坏情况下的保障,但实际攻击(尤其是成员推断)在中等隐私预算和辅助信息下仍更易针对这些异常值成功。本文提出风险均等化差分隐私合成框架,通过减少高风险记录对生成器的影响来优先保护它们。该机制分两阶段:第一阶段用小隐私预算估计每条记录的“异常度”;第二阶段在差分隐私学习中将记录权重设为其风险得分的反比。在高斯机制下,记录的隐私损失与其对输出的影响成正比,因此主动缩小异常值的贡献可获得更紧的个体级隐私边界。我们通过组合定理证明了端到端差分隐私保证,并为合成阶段推导出闭式个体级边界(评分阶段增加统一的每记录项)。模拟数据实验表明,风险加权显著降低了对高异常度记录的成员推断成功率;消融实验证明目标性降权是关键。在真实数据集(乳腺癌、成人、德国信贷)上,增益因数据集而异,凸显评分质量与合成流程之间的相互作用。
原文摘要 · Abstract (English)
When synthetic data is released, some individuals are harder to protect than others. A patient with a rare disease combination or a transaction with unusual characteristics stands out from the crowd. Differential privacy provides worst-case guarantees, but empirical attacks -- particularly membership inference -- succeed far more often against such outliers, especially under moderate privacy budgets and with auxiliary information. This paper introduces risk-equalized DP synthesis, a framework that prioritizes protection for high-risk records by reducing their influence on the learned generator. The mechanism operates in two stages: first, a small privacy budget estimates each record's "outlierness"; second, a DP learning procedure weights each record inversely to its risk score. Under Gaussian mechanisms, a record's privacy loss is proportional to its influence on the output -- so deliberately shrinking outliers' contributions yields tighter per-instance privacy bounds for precisely those records that need them most. We prove end-to-end DP guarantees via composition and derive closed-form per-record bounds for the synthesis stage (the scoring stage adds a uniform per-record term). Experiments on simulated data with controlled outlier injection show that risk-weighting substantially reduces membership inference success against high-outlierness records; ablations confirm that targeting -- not random downweighting -- drives the improvement. On real-world benchmarks (Breast Cancer, Adult, German Credit), gains are dataset-dependent, highlighting the interplay between scorer quality and synthesis pipeline.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。