arXiv:2411.16121stat.MLcs.LG2024-11被引 1

通过随机混合法生成隐私保护数据集,兼顾更高安全与更好实用性。

DP-CDA: An Algorithm for Enhanced Privacy Preservation in Dataset Synthesis Through Randomized Mixing

  • 按类别随机混合敏感数据,并注入可控随机性
  • 在相同隐私要求下,模型预测准确率显著提升
  • 适合需要强隐私保护的医疗、金融等高敏感领域

近年来,医疗、安全、金融和教育等领域数据量激增,为分析与决策带来机遇,但数据中常含敏感个人信息,引发严重隐私担忧。已有研究证明,即使数据匿名化,个体身份仍可通过信息模式被识别。现有机器学习与数据发布算法在处理高维数据时,难以平衡计算效率与隐私保护。为此,本文提出一种新型数据发布算法 DP-CDA,通过类特定的随机混合敏感数据并引入精确调控的随机性,实现形式化隐私保障。全面的隐私会计分析表明,相比现有方法,DP-CDA 提供更强隐私保护,在保持严格隐私水平的同时提升数据效用。实验评估显示,基于合成数据训练的预测模型具有更高准确性,且存在最优混合顺序可平衡隐私-效用权衡。结果表明,相同隐私约束下,DP-CDA 生成的数据集在实用性上优于传统算法。

原文摘要 · Abstract (English)

In recent years, the growth of data across various sectors, including healthcare, security, finance, and education, has created significant opportunities for analysis and informed decision-making. However, these datasets often contain sensitive and personal information, which raises serious privacy concerns. It has been shown in multiple works that a person's identity is intertwined with their data, even if the data is anonymized. Due to this lack of separation between a person's identity and their information, the patterns associated with an individual's information can uniquely identify them. Protecting individual privacy is crucial, yet many existing machine learning and data publishing algorithms struggle with high-dimensional data, facing challenges related to the trade-off between computational efficiency and privacy. To address these challenges, we introduce an effective data publishing algorithm \emph{DP-CDA}. Our proposed algorithm generates synthetic data by randomly mixing the privacy-sensitive data in a class-specific manner and inducing carefully tuned randomness to ensure formal privacy guarantees. Our comprehensive privacy accounting shows that the proposed DP-CDA provides a stronger privacy guarantee compared to existing methods, allowing for better utility while maintaining a stricter level of privacy. To evaluate the effectiveness of DP-CDA, we examine the accuracy of predictive models trained on the synthetic data, which serves as a measure of dataset utility. Importantly, we identify an optimal order of mixing that balances privacy-utility trade-off. Our results indicate that synthetic datasets produced using the DP-CDA can achieve superior utility compared to those generated by conventional data publishing algorithms, even when subject to the same privacy requirements.

隐私保护数据合成差分隐私数据安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。