arXiv:2504.08254cs.CRcs.LG2025-04中稿 · the Synthetic Data…被引 2

数据域提取方式影响合成数据隐私,不当方法会暴露敏感信息。

Understanding the Impact of Data Domain Extraction on Synthetic Data Privacy

  • 对比三种数据域提取策略:外部提供、直接从输入数据提取、差分隐私提取。
  • 直接提取数据域会破坏端到端差分隐私,使模型易受成员推断攻击。
  • 使用差分隐私提取或外部代表性数据域可有效防御主流隐私攻击。

隐私攻击,尤其是成员推断攻击(MIAs),被广泛用于评估表格型合成数据生成模型的隐私性,包括具有差分隐私(DP)保障的模型。这些攻击常利用异常值,因其位于数据域边界(如最小值和最大值)而尤为脆弱。然而,生成模型中数据域提取的作用及其对隐私攻击的影响尚未得到充分关注。本文考察了三种定义数据域的策略:假设其由外部提供(理想情况下来自公开数据)、直接从输入数据中提取,以及使用差分隐私机制提取。尽管第二种方法在流行实现和库中常见,但本文表明,它会破坏端到端差分隐私保证,使模型面临风险。相比之下,若数据域具有代表性,外部提供更为优选;而使用差分隐私提取也能有效抵御主流成员推断攻击,即使在高隐私预算下依然有效。

原文摘要 · Abstract (English)

Privacy attacks, particularly membership inference attacks (MIAs), are widely used to assess the privacy of generative models for tabular synthetic data, including those with Differential Privacy (DP) guarantees. These attacks often exploit outliers, which are especially vulnerable due to their position at the boundaries of the data domain (e.g., at the minimum and maximum values). However, the role of data domain extraction in generative models and its impact on privacy attacks have been overlooked. In this paper, we examine three strategies for defining the data domain: assuming it is externally provided (ideally from public data), extracting it directly from the input data, and extracting it with DP mechanisms. While common in popular implementations and libraries, we show that the second approach breaks end-to-end DP guarantees and leaves models vulnerable. While using a provided domain (if representative) is preferable, extracting it with DP can also defend against popular MIAs, even at high privacy budgets.

隐私保护差分隐私合成数据成员推断攻击

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。