用SMOTE增强差分隐私生成数据,兼顾隐私与实用性能。
SMOTE-DP: Improving Privacy-Utility Tradeoff with Synthetic Data
- 结合SMOTE与差分隐私机制生成合成数据
- 在保持高实用性的同时实现强隐私保护
- 适合注重隐私安全的数据发布场景
隐私保护的数据发布,包括合成数据共享,通常面临隐私与效用之间的权衡。合成数据相较于数据匿名化,在平衡这一权衡方面更有效,但仍存在挑战。由源数据训练的生成模型所产生的合成数据可能无意中泄露异常值信息。专门用于保护隐私的技术,如引入噪声以满足差分隐私,常导致难以预测且显著的效用损失。本文表明,通过合适的合成数据生成机制,可在不造成显著效用损失的情况下实现强隐私保护。利用产生收缩数据模式的合成数据生成器(如合成少数类过采样技术,SMOTE),可增强差分隐私数据生成器,融合两者优势。理论证明并实证显示,该SMOTE-DP方法能生成既确保稳健隐私保护又在下游学习任务中保持高实用性的合成数据。
原文摘要 · Abstract (English)
Privacy-preserving data publication, including synthetic data sharing, often experiences trade-offs between privacy and utility. Synthetic data is generally more effective than data anonymization in balancing this trade-off, however, not without its own challenges. Synthetic data produced by generative models trained on source data may inadvertently reveal information about outliers. Techniques specifically designed for preserving privacy, such as introducing noise to satisfy differential privacy, often incur unpredictable and significant losses in utility. In this work we show that, with the right mechanism of synthetic data generation, we can achieve strong privacy protection without significant utility loss. Synthetic data generators producing contracting data patterns, such as Synthetic Minority Over-sampling Technique (SMOTE), can enhance a differentially private data generator, leveraging the strengths of both. We prove in theory and through empirical demonstration that this SMOTE-DP technique can produce synthetic data that not only ensures robust privacy protection but maintains utility in downstream learning tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。