arXiv:2606.07640cs.CVcs.AI2026-06

在数据稀缺时,合成图像的保真度、隐私与实用性存在权衡。

No Free Lunch for Synthetic Images under Data Scarcity Conditions

论文配图:No Free Lunch for Synthetic Images under Data Scarcity Conditions
图 1 · 摘自论文原文
  • 构建多维度评估框架,同时衡量合成数据的保真度、隐私和实用价值
  • 隐私约束下,GAN与DDPM保持更高图像质量与下游任务性能
  • 适用于医疗影像等敏感领域中生成模型的选型与评估

本研究探讨了在数据稀缺和隐私敏感条件下,合成数据生成中保真度、隐私与实用性之间的权衡。我们提出一个联合评估框架,对三种主流生成模型(VAE、GAN、DDPM)在三个图像数据集(MNIST、OCTMNIST、OrganAMNIST)上的表现进行评估,涵盖通用与医学影像领域。当引入差分隐私机制训练时,三类模型行为差异显著:GAN与DDPM表现出更强鲁棒性,在不同噪声水平下仍维持较高保真度与下游任务效用;而VAE随隐私约束增强,性能下降更迅速。研究强调了对深度生成模型进行多维度评估的重要性,并指出隐私技术的应用会显著改变模型行为。

原文摘要 · Abstract (English)

This study investigates the trade-offs between fidelity, privacy, and utility in synthetic data generation under conditions of data scarcity and privacy sensitivity. We propose an evaluation framework that jointly assesses these three dimensions and apply it to three widely used generative models, VAE, GAN, and DDPM. The evaluation spans three image datasets, MNIST, OCTMNIST, and OrganAMNIST, encompassing both general-purpose and medical imaging domains. Notable differences arise between the three models in their behaviour when differential privacy mechanisms are introduced during training. GAN and DDPM demonstrate greater robustness, maintaining higher fidelity and downstream utility across a range of noise levels, while VAE degrades more rapidly as privacy constraints increase. This study highlights the importance of a multidimensional evaluation of deep generative models, also noting that their behaviour significantly differs when privacy techniques are applied.

生成模型隐私保护数据合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。