arXiv:2410.05506cs.CRcs.LG2024-10中稿 · 3rd IEEE Conferenc…被引 15

基于边缘概率的合成数据存在隐私漏洞,可被高效恢复个体信息。

Privacy Vulnerabilities in Marginals-based Synthetic Data

  • 提出新型成员推断攻击MAMA-MIA,利用算法类型信息加速破解
  • 在3种主流差分隐私合成算法上验证,攻击效率提升数量级
  • 揭示边缘概率类合成数据的隐蔽隐私风险,适合隐私安全研究者

作为隐私保护技术,合成数据生成(SDG)旨在保持与真实数据相似性的同时排除个人身份信息。许多SDG算法提供强大的差分隐私(DP)保障。然而,我们发现最强大的一类算法——即保留底层数据边际概率或类似统计量的算法——会泄露可被更高效恢复的个体信息。本文提出新型成员推断攻击MAMA-MIA,针对三种经典DP SDG算法(MST、PrivBayes、Private-GSD)进行评估。MAMA-MIA利用已知的SDG算法类型,能比现有领先攻击更准确、快数个数量级地获取隐藏数据信息。该方法最终赢得首届SNAKE(SaNitization Algorithm under attacK ... ε)竞赛。

原文摘要 · Abstract (English)

When acting as a privacy-enhancing technology, synthetic data generation (SDG) aims to maintain a resemblance to the real data while excluding personally-identifiable information. Many SDG algorithms provide robust differential privacy (DP) guarantees to this end. However, we show that the strongest class of SDG algorithms--those that preserve \textit{marginal probabilities}, or similar statistics, from the underlying data--leak information about individuals that can be recovered more efficiently than previously understood. We demonstrate this by presenting a novel membership inference attack, MAMA-MIA, and evaluate it against three seminal DP SDG algorithms: MST, PrivBayes, and Private-GSD. MAMA-MIA leverages knowledge of which SDG algorithm was used, allowing it to learn information about the hidden data more accurately, and orders-of-magnitude faster, than other leading attacks. We use MAMA-MIA to lend insight into existing SDG vulnerabilities. Our approach went on to win the first SNAKE (SaNitization Algorithm under attacK ... $\varepsilon$) competition.

隐私安全合成数据差分隐私成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。