用对比学习提升表格数据隐私评估,更准且易实现。
Contrastive Learning-Based privacy metrics in Tabular Synthetic Datasets
- 用对比学习嵌入数据,解决多类型属性带来的评估难题
- 结合相似度与攻击实验,验证了新方法与高级指标效果相当
- 适合关注数据隐私合规(如GDPR)的从业者快速评估
合成数据作为医疗、金融等领域的隐私增强技术受到关注。实际应用中需提供保护保障。现有两类方法:一类基于相似性,衡量训练数据与合成数据的相似程度;另一类基于攻击,通过攻击成功率评估安全性。本文提出一种基于对比学习的隐私评估方法,将数据嵌入更具代表性的空间,克服多种数据类型和属性带来的挑战,使直观的距离度量可用于相似性分析和攻击路径设计。在多个公开数据集上的实验表明,该方法在相似性与攻击评估中表现良好,其性能与明确建模GDPR隐私条件的复杂指标相当,且实现更高效、简便。
原文摘要 · Abstract (English)
Synthetic data has garnered attention as a Privacy Enhancing Technology (PET) in sectors such as healthcare and finance. When using synthetic data in practical applications, it is important to provide protection guarantees. In the literature, two family of approaches are proposed for tabular data: on the one hand, Similarity-based methods aim at finding the level of similarity between training and synthetic data. Indeed, a privacy breach can occur if the generated data is consistently too similar or even identical to the train data. On the other hand, Attack-based methods conduce deliberate attacks on synthetic datasets. The success rates of these attacks reveal how secure the synthetic datasets are. In this paper, we introduce a contrastive method that improves privacy assessment of synthetic datasets by embedding the data in a more representative space. This overcomes obstacles surrounding the multitude of data types and attributes. It also makes the use of intuitive distance metrics possible for similarity measurements and as an attack vector. In a series of experiments with publicly available datasets, we compare the performances of similarity-based and attack-based methods, both with and without use of the contrastive learning-based embeddings. Our results show that relatively efficient, easy to implement privacy metrics can perform equally well as more advanced metrics explicitly modeling conditions for privacy referred to by the GDPR.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。