对比合成图像在真实数据替代中的表现与隐私风险
SoK: Can Synthetic Images Replace Real Data? A Survey of Utility and Privacy of Synthetic Image Generation
- 按生成-采样-分类流程梳理合成图像方法与隐私攻防
- 用模型无关成员推理攻击评估隐私风险,验证合成数据有效性
- 给出不同场景下最优生成模型与发布策略建议
生成模型的发展推动了隐私保护数据合成(PPDS)领域进步,但该领域缺乏对各类合成图像生成方法在多样环境下的系统性综述与比较。本文针对以训练分类器为目的的合成图像生成,构建了从私有数据输入到最终分类器输出的生成-采样-分类全流程框架。系统分类现有图像生成方法、隐私攻击手段及缓解措施,并建立基准测试平台,采用模型无关的成员推理攻击(MIAs)量化隐私风险。通过系统评估多种合成方法,回答关键问题:合成数据能否有效替代真实数据?何种发布策略平衡效用与隐私?缓解措施是否改善效用-隐私权衡?哪些生成模型在不同场景表现最优?研究为实际应用中合成数据释放策略提供可操作的洞见。
原文摘要 · Abstract (English)
Advances in generative models have transformed the field of synthetic image generation for privacy-preserving data synthesis (PPDS). However, the field lacks a comprehensive survey and comparison of synthetic image generation methods across diverse settings. In particular, when we generate synthetic images for the purpose of training a classifier, there is a pipeline of generation-sampling-classification which takes private training as input and outputs the final classifier of interest. In this survey, we systematically categorize existing image synthesis methods, privacy attacks, and mitigations along this generation-sampling-classification pipeline. To empirically compare diverse synthesis approaches, we provide a benchmark with representative generative methods and use model-agnostic membership inference attacks (MIAs) as a measure of privacy risk. Through this study, we seek to answer critical questions in PPDS: Can synthetic data effectively replace real data? Which release strategy balances utility and privacy? Do mitigations improve the utility-privacy tradeoff? Which generative models perform best across different scenarios? With a systematic evaluation of diverse methods, our study provides actionable insights into the utility-privacy tradeoffs of synthetic data generation methods and guides the decision on optimal data releasing strategies for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。