通过分析模型对图像和文本的去噪差异,可有效检测图像生成模型训练数据中的成员信息。
Black-box Membership Inference Attacks on the Pre-training Data of Image-generation Models

- 利用跨模态扰动设计黑盒攻击,通过图文去噪差异识别成员数据。
- 在公开与自建数据集上均超越已有方法,包括能访问内部特征的基线。
- 适用于无法获取模型内部信息的闭源平台,具有实际应用价值。
基于扩散的图像生成模型快速发展,引发对人类创作数据版权与隐私泄露的担忧。成员推理攻击(MIAs)已成为检测模型训练中未经授权数据使用的重要工具。现有方法通常以模型对扰动样本的去噪能力作为成员身份指标,但该特征依赖于模型记忆程度,在较少暴露的数据(如预训练数据)上性能显著下降。尽管部分方法尝试利用模型内部特征提升检测效果,但这些特征在主流闭源平台中通常不可访问,限制了其实用性。本文提出一种黑盒成员推理攻击框架(SD-MIA),通过分析扩散模型对目标图像及其对应扰动文本指令的去噪行为,揭示更显著的成员线索。我们基于跨模态数据扰动机制构建攻击方法,并在公开基准数据集与新构建的数据集上进行广泛实验,每组包含分布相同的成员与非成员样本。结果表明,SD-MIA在性能上优于现有基线,甚至超越那些可访问内部特征的不公平对比方法。
原文摘要 · Abstract (English)
The rapid advancement of diffusion-based image generation models has raised serious concerns regarding potential copyright and privacy infringements involving human-created data. Membership inference attacks (MIAs) have emerged as a promising tool for identifying unauthorized data usage during model training. Existing methods typically assess the ability of model to denoise perturbed suspect images as an indicator of membership status. However, the discriminative power of such features is highly dependent on the degree of model memorization and deteriorates significantly when applied to less exposed data (e.g., pre-training data). Although several methods attempt to enhance detection by leveraging internal model features, these features are generally inaccessible in mainstream closed-source image generation platforms, limiting their practicality. In this paper, we demonstrate that analyzing how a black-box diffusion model denoises a target image and corresponding perturbed textual instructions can reveal more distinctive membership cues. Based on this insight, we propose a black-box membership inference attack framework (named SD-MIA) that leverages a cross-modal data perturbation mechanism to detect pre-training data in diffusion models. We conduct extensive experiments on both a public benchmark dataset and a newly constructed dataset, each comprising pre-training membership and non-membership samples with identical distributions. Experimental results demonstrate that SD-MIA achieves superior performance compared to existing baselines, including those with the unfair advantage of accessing internal model features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。