合成人脸数据可能泄露真实训练数据,存在隐私风险。
Have I Seen You? Embedding Behavior Signals Synthetic Face Dataset Membership

- 通过数据集级成员推理攻击,识别训练用的合成数据集。
- 100%准确恢复合成数据集,54.5%识别生成器的真实来源数据集。
- 提醒研究者和开发者:合成数据不能完全规避隐私泄露。
合成人脸数据集在生物特征识别中被广泛用于降低隐私暴露和数据访问限制。然而,生成这些数据的模型是基于真实人脸训练的,因此合成数据可能仍保留其真实来源的痕迹。本文提出一种数据集级别的成员推理攻击,首先确定用于训练人脸识别模型的合成数据集,再推断生成该合成数据的原始真实数据集。在11个面部识别模型、11个合成数据集和7个真实数据集的实验中,该攻击在100%情况下成功恢复合成训练数据集,并在54.5%的情况下识别出生成器的真实来源数据集。结果表明,合成数据可能仍携带真实训练数据的层级痕迹,因此在部署时需更强的泄漏缓解措施。
原文摘要 · Abstract (English)
Synthetic face datasets are increasingly used to reduce privacy exposure and data access constraints in biometric recognition. Yet the generators that produce these datasets are trained on real faces, so synthetic data may still reveal their real source data. We study this risk through a dataset-level membership inference attack that first identifies the synthetic dataset used to train a face recognizer and then infers the real dataset used to train the generator. Across 11 face recognition models, 11 synthetic datasets, and 7 real datasets, the attack recovers the synthetic training dataset in 100% of cases and identifies the generator's source dataset in 54.5% of cases. These results show that synthetic data can retain dataset-level traces of real training data and that privacy-preserving deployment requires stronger leakage mitigation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。