arXiv:2503.14681cs.CRcs.AI2025-03中稿 · CCS 2025被引 18

构建统一基准评估差分隐私图像生成,纠正错误评估方法。

DPImageBench: A Unified Benchmark for Differentially Private Image Synthesis

  • 设计多维度评估框架,涵盖11种方法、9个数据集和7项指标。
  • 发现预训练数据分布相似性影响生成效果,非越通用越好。
  • 低维特征加噪比高维特征更抗隐私预算削减,适合低隐私场景。

差分隐私图像生成旨在生成保留敏感图像特性的合成图像,同时保护数据集中个体图像的隐私。尽管近期取得进展,我们发现各研究采用的评估协议不一致甚至存在缺陷,这不仅阻碍对现有方法的理解,也制约未来发展。为此,本文提出DPImageBench,用于差分隐私图像生成的统一基准。其设计包含三个维度:(1)方法层面,系统分析11种主流方法在模型架构、预训练策略与隐私机制上的差异;(2)评估层面,涵盖9个数据集与7项保真度与效用指标,特别指出基于敏感测试集最高准确率选择下游分类器的做法违反差分隐私且高估性能,而本基准已纠正此问题;(3)平台层面,提供标准化接口,兼容当前及未来实现。通过该基准,我们发现:预训练于公开图像数据集未必有益,关键在于预训练与敏感图像的分布相似性;在低隐私预算下,对低维特征(如高层语义)加噪的方法优于对高维特征(如权重梯度)加噪,表现更优。

原文摘要 · Abstract (English)

Differentially private (DP) image synthesis aims to generate artificial images that retain the properties of sensitive images while protecting the privacy of individual images within the dataset. Despite recent advancements, we find that inconsistent--and sometimes flawed--evaluation protocols have been applied across studies. This not only impedes the understanding of current methods but also hinders future advancements. To address the issue, this paper introduces DPImageBench for DP image synthesis, with thoughtful design across several dimensions: (1) Methods. We study eleven prominent methods and systematically characterize each based on model architecture, pretraining strategy, and privacy mechanism. (2) Evaluation. We include nine datasets and seven fidelity and utility metrics to thoroughly assess them. Notably, we find that a common practice of selecting downstream classifiers based on the highest accuracy on the sensitive test set not only violates DP but also overestimates the utility scores. DPImageBench corrects for these mistakes. (3) Platform. Despite the methods and evaluation protocols, DPImageBench provides a standardized interface that accommodates current and future implementations within a unified framework. With DPImageBench, we have several noteworthy findings. For example, contrary to the common wisdom that pretraining on public image datasets is usually beneficial, we find that the distributional similarity between pretraining and sensitive images significantly impacts the performance of the synthetic images and does not always yield improvements. In addition, adding noise to low-dimensional features, such as the high-level characteristics of sensitive images, is less affected by the privacy budget compared to adding noise to high-dimensional features, like weight gradients. The former methods perform better than the latter under a low privacy budget.

差分隐私图像生成评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。