首个针对大模型的成员推断攻击基准,可检测训练数据中敏感图像。
Membership Inference Attacks against Large Vision-Language Models
- 构建首个面向视觉语言模型的成员推断攻击基准
- 提出基于令牌级置信度的图像数据检测新方法
- 设计适用于图文数据的新型评估指标 MaxRényi-K%
大型视觉语言模型(VLLMs)在多模态任务中展现出强大能力,但其训练数据可能包含私人照片、医疗记录等敏感信息,引发数据安全问题。由于缺乏标准化数据集和有效方法,检测不当使用数据仍是一个未解难题。本文首次为各类VLLMs构建了成员推断攻击(MIA)基准,提出专用于令牌级图像检测的MIA新流程,并引入基于模型输出置信度的新指标MaxRényi-K%,适用于文本与图像数据。本研究有助于深化对VLLMs中成员推断攻击的理解与方法发展。代码与数据集已公开于https://github.com/LIONS-EPFL/VL-MIA。
原文摘要 · Abstract (English)
Large vision-language models (VLLMs) exhibit promising capabilities for processing multi-modal tasks across various application scenarios. However, their emergence also raises significant data security concerns, given the potential inclusion of sensitive information, such as private photos and medical records, in their training datasets. Detecting inappropriately used data in VLLMs remains a critical and unresolved issue, mainly due to the lack of standardized datasets and suitable methodologies. In this study, we introduce the first membership inference attack (MIA) benchmark tailored for various VLLMs to facilitate training data detection. Then, we propose a novel MIA pipeline specifically designed for token-level image detection. Lastly, we present a new metric called MaxRényi-K%, which is based on the confidence of the model output and applies to both text and image data. We believe that our work can deepen the understanding and methodology of MIAs in the context of VLLMs. Our code and datasets are available at https://github.com/LIONS-EPFL/VL-MIA.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。