数据蒸馏反成隐私泄露漏洞,合成数据暗藏真实数据痕迹
Turning Black Box into White Box: Dataset Distillation Leaks
- 用合成数据压缩真实数据时,模型训练轨迹被隐含保留
- 攻击可精准还原蒸馏算法与模型结构,甚至复原敏感样本
- 适合关注数据隐私安全的研究者和工业界应用者
数据蒸馏将大规模真实数据集压缩为小型合成数据集,使在合成数据上训练的模型性能可媲美真实数据训练的结果。尽管合成数据被视为具有隐私保护性,但本文揭示现有蒸馏方法会导致严重隐私泄露:合成数据隐式编码了所蒸馏模型的权重轨迹,变得过度信息密集且易被攻击者利用。为此,我们提出信息泄露攻击(IRA),针对当前最先进的蒸馏技术进行测试。实验表明,IRA能准确预测蒸馏算法与模型架构,并成功推断出原始数据中的成员身份,甚至恢复敏感样本。该研究警示了数据蒸馏在实际应用中的潜在风险。
原文摘要 · Abstract (English)
Dataset distillation compresses a large real dataset into a small synthetic one, enabling models trained on the synthetic data to achieve performance comparable to those trained on the real data. Although synthetic datasets are assumed to be privacy-preserving, we show that existing distillation methods can cause severe privacy leakage because synthetic datasets implicitly encode the weight trajectories of the distilled model, they become over-informative and exploitable by adversaries. To expose this risk, we introduce the Information Revelation Attack (IRA) against state-of-the-art distillation techniques. Experiments show that IRA accurately predicts both the distillation algorithm and model architecture, and can successfully infer membership and recover sensitive samples from the real dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。