首次系统评估视觉语言模型遗忘能力,发现现有方法易被重新激活知识。
On the Robustness of Machine Unlearning for Vision-Language Models

- 提出三种攻击范式,检验遗忘知识是否可通过提示或微调恢复。
- 实验表明多数方法仅隐藏而非彻底删除目标知识,存在安全隐患。
- 适合关注多模态模型隐私与安全的研究者和工程师参考。
视觉语言模型(VLMs)可能从训练数据中记忆不当信息,促使机器遗忘技术日益受到关注。本文首次对VLM遗忘进行系统性调研与鲁棒性分析,提供现有方法的全面分类,并在多种提示设置下统一评估。随后,提出三种攻击范式,检验遗忘的多模态知识是否可通过上下文提示或下游重训练重新激活。大量实验表明,许多现有方法在这些攻击下仍显脆弱,说明当前策略常为隐藏而非完全移除目标知识。本研究揭示了现有VLM遗忘方法的鲁棒性局限,强调需发展更可靠的多模态遗忘机制。代码已开源:https://github.com/XMUDeepLIT/VLM-UnL-Attack。
原文摘要 · Abstract (English)
Vision-language models (VLMs) may memorize undesirable information from training data, motivating growing interest in machine unlearning. In this work, we present the first systematic survey and robustness analysis of VLM unlearning. We provide a comprehensive taxonomy and review of existing VLM unlearning methods, together with unified evaluations under multiple prompt settings. We then propose three attack paradigms to examine whether forgotten multimodal knowledge can be reactivated through contextual prompting or downstream retraining. Extensive experiments show that many existing methods remain vulnerable under these attacks, indicating that current approaches often hide rather than fully remove target knowledge. Our study provides new insights into the robustness and limitations of current VLM unlearning methods and highlights the need for more reliable multimodal unlearning strategies. Code is available at https://github.com/XMUDeepLIT/VLM-UnL-Attack.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。