arXiv:2410.24015cs.CV2024-10中稿 · NeurIPS被引 9

发现合成人脸数据集会泄露训练生成模型的真实人脸信息。

Unveiling Synthetic Faces: How Synthetic Datasets Can Expose Real Identities

  • 设计简单有效的成员推理攻击,检测合成数据中的真实数据泄露。
  • 在6个主流合成人脸数据集中均发现原始真实样本被泄露。
  • 首次揭示生成模型训练数据对合成数据的隐私泄露问题,警示数据安全。

合成数据生成在计算机视觉应用中日益流行。现有先进的人脸识别模型多基于从互联网爬取的大规模人脸数据集训练,引发隐私与伦理担忧。为此,部分研究提出生成合成人脸数据集以替代真实数据。然而,这些方法依赖于基于真实人脸图像训练的生成模型。本文设计一种简单而有效的方法,系统性地研究现有合成人脸数据集是否泄露了生成模型的训练数据。我们在6个最先进的合成人脸数据集上进行了广泛实验,结果表明所有数据集中均存在原始真实数据样本的泄露。据我们所知,这是首个揭示生成模型训练数据向合成人脸数据集泄露的研究。本工作揭示了合成人脸数据集中的隐私风险,为未来负责任合成数据的生成提供了重要方向。

原文摘要 · Abstract (English)

Synthetic data generation is gaining increasing popularity in different computer vision applications. Existing state-of-the-art face recognition models are trained using large-scale face datasets, which are crawled from the Internet and raise privacy and ethical concerns. To address such concerns, several works have proposed generating synthetic face datasets to train face recognition models. However, these methods depend on generative models, which are trained on real face images. In this work, we design a simple yet effective membership inference attack to systematically study if any of the existing synthetic face recognition datasets leak any information from the real data used to train the generator model. We provide an extensive study on 6 state-of-the-art synthetic face recognition datasets, and show that in all these synthetic datasets, several samples from the original real dataset are leaked. To our knowledge, this paper is the first work which shows the leakage from training data of generator models into the generated synthetic face recognition datasets. Our study demonstrates privacy pitfalls in synthetic face recognition datasets and paves the way for future studies on generating responsible synthetic face datasets.

人脸识别合成数据隐私泄露

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。