arXiv:2605.17341cs.CVcs.AI2026-05

通过跨模态对齐检测视觉语言模型的训练数据成员身份,单样本黑盒攻击更有效。

Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment

论文配图:Single-Sample Black-Box Membership Inference Attack against Vision-Language Models via Cross-modal Semantic Alignment
图 1 · 摘自论文原文
  • 基于图像与标题在联合嵌入空间中的对齐强度差异设计攻击
  • 单样本下在VL-MIA/Flicker数据集上达到0.821 AUC
  • 适用于真实场景中无法获取内部信息的黑盒攻击

视觉语言模型(VLMs)虽取得显著进展,但其依赖大规模数据集及对训练数据的无意记忆带来严重数据安全风险。成员推断攻击(MIAs)旨在评估此类风险,判断某数据样本是否属于模型训练集。然而,现有针对VLM的MIAs存在关键瓶颈:灰盒方法依赖通常在实际API中受限的内部logits,而黑盒方法依赖大规模统计分布,在单样本场景下表现不佳。为此,我们从跨模态语义对齐角度研究MIAs,观察到成员图像因训练记忆表现出更强的图像-标题对齐,而非成员生成的标题则可能偏离原始视觉内容。基于此洞察,我们提出一种专为严格黑盒与单样本设置设计的新颖MIA框架,通过量化联合嵌入空间内的对齐程度,规避了不切实际的假设。我们在三个开源和两个闭源VLM上进行了广泛实验。在VL-MIA/Flicker数据集上,该方法对LLaVA-1.5实现0.821 AUC,显著优于现有基线。此外,其在多种图像扰动下仍保持鲁棒性,凸显其实用性。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) have achieved remarkable success, yet their reliance on massive datasets and unintended memorization of training data raise significant data security risk. Membership Inference Attacks (MIAs) aim to assess these risks by determining whether a data sample was included in a model's training set. However, existing MIA methods against VLMs face critical bottlenecks: gray-box method relies on internal logits that are typically restricted in real-world Application Programming Interfaces (APIs), while black-box method depends on large-scale statistical distributions, which struggle in single-sample scenarios. To this end, we investigate MIAs from the perspective of cross-modal semantic alignment, and observe that member images exhibit significantly stronger image-caption alignment due to training memorization, whereas generated captions for non-members may deviate from the original visual content. Leveraging this insight, we propose a novel MIA framework designed for strict black-box and single-sample setting that quantifies such alignment within a joint embedding space, thereby bypassing these unrealistic assumptions. We conducted extensive experiments on three open-source and two closed-source VLMs. On the VL-MIA/Flicker dataset, our method achieves an AUC of 0.821 against LLaVA-1.5, significantly outperforming existing baselines. Furthermore, it remains robust under diverse image perturbations, highlighting its practicality.

成员推断视觉语言模型黑盒攻击安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。