arXiv:2501.18624cs.CRcs.AI2025-01中稿 · USENIX'25被引 41

首次对视觉语言模型发起成员推断攻击,揭示训练数据泄露风险。

Membership Inference Attacks Against Vision-Language Models

  • 利用温度参数敏感性设计新型成员推断方法
  • 在仅5个样本上仍达0.8以上AUC性能
  • 适合关注AI隐私安全的研究者与开发者

视觉语言模型(VLMs)结合预训练视觉编码器与大语言模型,在多模态理解与对话方面表现出色,被视为下一代技术革命的催化剂。然而,当前研究多聚焦于提升多模态交互能力,对数据滥用与泄露风险的关注却严重不足。本文首次从成员推断攻击(MIA)视角,系统分析VLMs中的数据泄露问题,重点关注指令微调数据中可能包含的敏感或未经授权信息。针对现有MIA方法的局限性,提出一种基于样本集合及其对温度参数敏感性的新方法。据此设计四种不同背景知识水平的推断策略,覆盖从简单到最复杂场景。全面评估表明,该方法能精准判断数据成员身份,例如在仅5个样本的测试集上,针对LLaVA模型实现超过0.8的AUC值。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs), built on pre-trained vision encoders and large language models (LLMs), have shown exceptional multi-modal understanding and dialog capabilities, positioning them as catalysts for the next technological revolution. However, while most VLM research focuses on enhancing multi-modal interaction, the risks of data misuse and leakage have been largely unexplored. This prompts the need for a comprehensive investigation of such risks in VLMs. In this paper, we conduct the first analysis of misuse and leakage detection in VLMs through the lens of membership inference attack (MIA). In specific, we focus on the instruction tuning data of VLMs, which is more likely to contain sensitive or unauthorized information. To address the limitation of existing MIA methods, we introduce a novel approach that infers membership based on a set of samples and their sensitivity to temperature, a unique parameter in VLMs. Based on this, we propose four membership inference methods, each tailored to different levels of background knowledge, ultimately arriving at the most challenging scenario. Our comprehensive evaluations show that these methods can accurately determine membership status, e.g., achieving an AUC greater than 0.8 targeting a small set consisting of only 5 samples on LLaVA.

成员推断视觉语言模型隐私安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。