arXiv:2601.22434cs.CRcs.CY2026-01被引 2

从模型视角重新审视合成数据匿名性,揭示其隐私风险。

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

  • 以生成模型能力为视角,评估合成数据的隐私风险。
  • 指出仅靠合成数据无法确保匿名,需结合攻击场景分析。
  • 对比差分隐私与相似性度量,证明前者更有效保障隐私。

训练生成式机器学习模型以生成合成表格数据,已成为提升数据共享中隐私保护的流行方法。由于通常涉及处理敏感个人信息,发布训练好的模型或生成的合成数据集仍可能带来隐私风险。然而,当前研究、商业部署及像通用数据保护条例(GDPR)这样的隐私法规,大多在个体数据集层面评估匿名性。本文从模型中心视角重新思考合成数据的匿名性声明,主张有意义的评估必须考虑底层生成模型的能力和特性,并基于最新的隐私攻击进行。这一视角更贴近真实产品与部署场景,其中训练好的模型往往可被直接交互或查询。我们在此类访问假设下解读GDPR对个人数据和匿名化的定义,识别出必须缓解的可识别性风险,并将其映射到不同威胁环境下的隐私攻击。研究发现,仅依靠合成数据技术不足以实现充分匿名。最后,我们比较了常与合成数据搭配使用的两种机制——差分隐私(DP)和基于相似性的隐私度量(SBPMs),认为尽管DP能提供对可识别性风险的稳健防护,但SBPMs缺乏足够保障。总体而言,本工作将监管中的可识别性概念与模型中心的隐私攻击相连接,使研究人员、从业者和政策制定者能够更负责任、可信地评估合成数据系统。

原文摘要 · Abstract (English)

Training generative machine learning models to produce synthetic tabular data has become a popular approach for enhancing privacy in data sharing. As this typically involves processing sensitive personal information, releasing either the trained model or generated synthetic datasets can still pose privacy risks. Yet, recent research, commercial deployments, and privacy regulations like the General Data Protection Regulation (GDPR) largely assess anonymity at the level of an individual dataset. In this paper, we rethink anonymity claims about synthetic data from a model-centric perspective and argue that meaningful assessments must account for the capabilities and properties of the underlying generative model and be grounded in state-of-the-art privacy attacks. This perspective better reflects real-world products and deployments, where trained models are often readily accessible for interaction or querying. We interpret the GDPR's definitions of personal data and anonymization under such access assumptions to identify the types of identifiability risks that must be mitigated and map them to privacy attacks across different threat settings. We then argue that synthetic data techniques alone do not ensure sufficient anonymization. Finally, we compare the two mechanisms most commonly used alongside synthetic data -- Differential Privacy (DP) and Similarity-based Privacy Metrics (SBPMs) -- and argue that while DP can offer robust protections against identifiability risks, SBPMs lack adequate safeguards. Overall, our work connects regulatory notions of identifiability with model-centric privacy attacks, enabling more responsible and trustworthy regulatory assessment of synthetic data systems by researchers, practitioners, and policymakers.

合成数据隐私攻击差分隐私GDPR

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。