让多模态大模型精准删除特定视觉概念,不伤及其他内容。
AUVIC: Adversarial Unlearning of Visual Concepts for Multi-modal Large Language Models
- 用对抗扰动精准定位并清除目标视觉概念。
- 在保留其他概念性能的同时,实现最高遗忘率。
- 专为多模态模型设计首个视觉概念遗忘评测基准。
多模态大语言模型(MLLM)在大规模数据集上优化后表现优异,但这些数据常包含敏感或受版权保护的内容,引发数据隐私问题。监管要求的‘被遗忘权’推动机器遗忘技术发展,该技术可在不重新训练的前提下移除目标数据。尽管文本领域已有研究,视觉概念在MLLM中的遗忘仍待探索。主要挑战在于精确删除目标视觉概念,同时避免影响相关实体。为此,我们提出AUVIC,一种针对MLLM的新型视觉概念遗忘框架。AUVIC通过引入对抗扰动实现精准遗忘,有效隔离目标概念,避免对相似实体的意外影响。为评估方法,我们构建了VCUBench,这是首个用于评估组内视觉概念遗忘的基准。实验表明,AUVIC在实现顶尖遗忘率的同时,对非目标概念性能影响极小。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) achieve impressive performance once optimized on massive datasets. Such datasets often contain sensitive or copyrighted content, raising significant data privacy concerns. Regulatory frameworks mandating the 'right to be forgotten' drive the need for machine unlearning. This technique allows for the removal of target data without resource-consuming retraining. However, while well-studied for text, visual concept unlearning in MLLMs remains underexplored. A primary challenge is precisely removing a target visual concept without disrupting model performance on related entities. To address this, we introduce AUVIC, a novel visual concept unlearning framework for MLLMs. AUVIC applies adversarial perturbations to enable precise forgetting. This approach effectively isolates the target concept while avoiding unintended effects on similar entities. To evaluate our method, we construct VCUBench. It is the first benchmark designed to assess visual concept unlearning in group contexts. Experimental results demonstrate that AUVIC achieves state-of-the-art target forgetting rates while incurs minimal performance degradation on non-target concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。