测试扩散模型生成的表格数据能否抵御成员推断攻击
MIDST Challenge at SaTML 2025: Membership Inference over Diffusion-models-based Synthetic Tabular data
- 针对扩散模型生成的表格数据设计黑盒与白盒成员推断攻击方法
- 发现合成表格数据在多类型混合表和多关系表中均存在隐私泄露风险
- 为评估合成数据隐私性提供可复现的基准测试框架,适合隐私研究者
合成数据常被视为数据匿名化与隐私保护发布的一剂良方。基于扩散模型生成的合成数据,预期在保留原始数据统计特性的同时,具备对隐私攻击的鲁棒性。尽管扩散模型已在多种数据类型上取得进展,其在表格数据上的隐私韧性仍缺乏系统研究。MIDST挑战旨在量化评估扩散模型生成的表格数据在成员推断攻击(MIAs)下的隐私收益,特别关注其对复杂异构表格数据的防护能力。针对表格数据的多样性与复杂性,挑战探索了多种目标模型,包括用于混合数据类型的单张表格扩散模型,以及具有关联约束的多关系表格扩散模型。该挑战催生了新型黑盒与白盒成员推断攻击方法,专为这些目标扩散模型量身定制,从而实现对其隐私效用的全面评估。MIDST GitHub仓库已公开:https://github.com/VectorInstitute/MIDST
原文摘要 · Abstract (English)
Synthetic data is often perceived as a silver-bullet solution to data anonymization and privacy-preserving data publishing. Drawn from generative models like diffusion models, synthetic data is expected to preserve the statistical properties of the original dataset while remaining resilient to privacy attacks. Recent developments of diffusion models have been effective on a wide range of data types, but their privacy resilience, particularly for tabular formats, remains largely unexplored. MIDST challenge sought a quantitative evaluation of the privacy gain of synthetic tabular data generated by diffusion models, with a specific focus on its resistance to membership inference attacks (MIAs). Given the heterogeneity and complexity of tabular data, multiple target models were explored for MIAs, including diffusion models for single tables of mixed data types and multi-relational tables with interconnected constraints. MIDST inspired the development of novel black-box and white-box MIAs tailored to these target diffusion models as a key outcome, enabling a comprehensive evaluation of their privacy efficacy. The MIDST GitHub repository is available at https://github.com/VectorInstitute/MIDST
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。