arXiv:2503.04980cs.CRcs.AI2025-03被引 27

提出合成数据隐私评估框架,解决现有指标无法衡量身份泄露问题。

A Consensus Privacy Metrics Framework for Synthetic Data

  • 通过专家共识构建隐私评估框架,明确关键隐私风险类型。
  • 指出当前相似性指标无法检测身份泄露,不推荐使用。
  • 强调成员身份与属性泄露的重要性,适合数据安全研究者参考。

合成数据生成是共享个体级数据的一种方法。然而,为满足立法要求,必须证明个体隐私得到充分保护。目前尚无统一的合成数据隐私度量标准。通过专家小组和共识过程,我们建立了评估合成数据隐私的框架。研究发现,现有相似性指标无法有效衡量身份泄露,其使用应被避免;对于差分隐私合成数据,除接近零的隐私预算外,其他值均难以解释。专家一致认为成员身份泄露和属性泄露至关重要,二者均涉及在不直接揭示身份的情况下推断个人敏感信息。该框架提供了针对此类泄露的有效度量建议,并指出了未来研究可推动合成数据广泛采用的具体方向。

原文摘要 · Abstract (English)

Synthetic data generation is one approach for sharing individual-level data. However, to meet legislative requirements, it is necessary to demonstrate that the individuals' privacy is adequately protected. There is no consolidated standard for measuring privacy in synthetic data. Through an expert panel and consensus process, we developed a framework for evaluating privacy in synthetic data. Our findings indicate that current similarity metrics fail to measure identity disclosure, and their use is discouraged. For differentially private synthetic data, a privacy budget other than close to zero was not considered interpretable. There was consensus on the importance of membership and attribute disclosure, both of which involve inferring personal information about an individual without necessarily revealing their identity. The resultant framework provides precise recommendations for metrics that address these types of disclosures effectively. Our findings further present specific opportunities for future research that can help with widespread adoption of synthetic data.

合成数据隐私评估差分隐私

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。