首个跨模态通用隐私攻击框架,可统一检测文本与图像生成模型的训练数据成员身份。
One Framework for All: Cross-Modal Membership Inference for Generative Models

- 利用生成输出分布近似训练数据分布的共性,构建跨模态统一推理框架。
- 在黑盒环境下对三种生成模型均实现更高精度的成员推断,优于单模态专用方法。
- 适用于微调和预训练数据的攻击,适合研究生成模型隐私漏洞的学者使用。
文本到文本、文本到图像及图像到文本三类大型生成模型均存在显著隐私风险。其中,成员推断攻击(MIA)旨在判断某数据点是否曾被用于模型训练。尽管已有研究针对这三类模型开展过MIA分析,但现有方法彼此孤立且不可跨模态应用,限制了实际价值。为此,本文首次提出一个统一的跨模态成员推断框架,覆盖上述三类生成模型。核心思想基于一个模态无关的观察:生成模型的输出分布可近似其训练数据分布。我们在此基础上,在共享嵌入空间中建模生成输出与外部非成员样本的分布,并通过似然比检验进行成员推断。我们在严格黑盒设置下,分别在部分知识与零知识威胁模型中进行了广泛实验,评估了对微调与预训练数据的攻击效果。结果表明,本方法在各项指标上均显著优于现有最优单模态方法。
原文摘要 · Abstract (English)
Large generative models across text-to-text, text-to-image, and image-to-text modalities have been shown to pose significant privacy risks. One fundamental threat is membership inference attacks (MIA), which aim to determine whether a given data point was used in a model's training set. Although prior work has investigated MIAs against these three classes of generative models, existing approaches treat them in isolation and are not cross-applicable, thereby limiting their real-world utility. To address this limitation, we present the first comprehensive study of a unified membership inference framework that applies across text-to-text, text-to-image, and image-to-text modalities. Our approach is grounded in a key modality-agnostic observation: the output distribution of a generative model can approximate its training data distribution. Leveraging this property, we model the distributions of model-generated outputs and auxiliary non-member samples in a shared embedding space, and perform membership inference via likelihood ratio testing. We conduct extensive experiments in a strict black-box setting under both partial-knowledge and zero-knowledge threat models, and evaluate membership inference against both fine-tuning and pre-training data. Experimental results demonstrate our approach's superior performance in comparison to existing state-of-the-art methods, which are typically optimized for a single model class.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。