arXiv:2505.01456cs.CLcs.AI2025-05被引 17

构建多模态遗忘基准,评估大模型删除敏感信息的能力。

Unlearning Sensitive Information in Multimodal LLMs: Benchmark and Attack-Defense Evaluation

  • 提出新基准UnLOK-VQA,生成带语义距离的图文对测试遗忘效果。
  • 多模态攻击比单模态更有效,大模型在编辑后更具鲁棒性。
  • 揭示模型内部状态是防御关键,适合安全与隐私研究者参考。

基于海量数据训练的大型语言模型可能无意中习得个人敏感信息和有害内容,这一风险在多模态大模型(MLLMs)中尤为突出,因其融合图像与文本信息。攻击者可通过多模态提示提取敏感内容。为评估模型针对性遗忘能力,需高质量、标注精细的图文对。现有研究集中于文本遗忘,多模态遗忘仍待探索。为此,我们构建了多模态遗忘基准UnLOK-VQA(Unlearning Outside Knowledge VQA),并设计攻击-防御框架评估删除特定多模态知识的方法。通过自动化流水线生成不同语义距离样本,并经人工筛选确保质量。评估六种防御目标对抗七种攻击(四白盒、三黑盒),包括一种利用隐藏状态可解释性的新型白盒方法。结果表明:多模态攻击性能优于单模态;最有效的防御策略是从模型内部状态移除答案信息;更大模型表现出更强的后期编辑鲁棒性,暗示规模提升安全性。UnLOK-VQA为推进多模态大模型遗忘研究提供严格基准。

原文摘要 · Abstract (English)

LLMs trained on massive datasets may inadvertently acquire sensitive information such as personal details and potentially harmful content. This risk is further heightened in multimodal LLMs as they integrate information from multiple modalities (image and text). Adversaries can exploit this knowledge through multimodal prompts to extract sensitive details. Evaluating how effectively MLLMs can forget such information (targeted unlearning) necessitates the creation of high-quality, well-annotated image-text pairs. While prior work on unlearning has focused on text, multimodal unlearning remains underexplored. To address this gap, we first introduce a multimodal unlearning benchmark, UnLOK-VQA (Unlearning Outside Knowledge VQA), as well as an attack-and-defense framework to evaluate methods for deleting specific multimodal knowledge from MLLMs. We extend a visual question-answering dataset using an automated pipeline that generates varying-proximity samples for testing generalization and specificity, followed by manual filtering for maintaining high quality. We then evaluate six defense objectives against seven attacks (four whitebox, three blackbox), including a novel whitebox method leveraging interpretability of hidden states. Our results show multimodal attacks outperform text- or image-only ones, and that the most effective defense removes answer information from internal model states. Additionally, larger models exhibit greater post-editing robustness, suggesting that scale enhances safety. UnLOK-VQA provides a rigorous benchmark for advancing unlearning in MLLMs.

多模态遗忘学习模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。