arXiv:2502.12976cs.CRcs.LG2025-02ICLR被引 15

四种合成数据训练方法隐私保护效果差异大,不能盲目信赖。

Does Training with Synthetic Data Truly Protect Privacy?

  • 对比四种合成数据训练方法的隐私表现
  • 实证发现不同方法隐私保护能力悬殊
  • 提醒需严格评估才能避免虚假安全

随着合成数据在机器学习中的广泛应用,许多无需正式差分隐私保证的方法使用合成数据进行训练,并常声称能保护原始数据隐私。本文研究了四种不同训练范式:核心集选择、数据集蒸馏、无数据知识蒸馏以及扩散模型生成的合成数据。尽管均采用合成数据训练,但它们在隐私保护方面的结论截然不同。研究警示,仅依赖经验方法评估数据隐私存在风险,必须进行严谨评估,否则可能产生虚假的隐私保护错觉。

原文摘要 · Abstract (English)

As synthetic data becomes increasingly popular in machine learning tasks, numerous methods--without formal differential privacy guarantees--use synthetic data for training. These methods often claim, either explicitly or implicitly, to protect the privacy of the original training data. In this work, we explore four different training paradigms: coreset selection, dataset distillation, data-free knowledge distillation, and synthetic data generated from diffusion models. While all these methods utilize synthetic data for training, they lead to vastly different conclusions regarding privacy preservation. We caution that empirical approaches to preserving data privacy require careful and rigorous evaluation; otherwise, they risk providing a false sense of privacy.

合成数据隐私保护机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。