arXiv:2512.11325cs.CVcs.AI2025-12

通过视觉知识蒸馏实现多模态模型的精准删减。

Robust MLLM Unlearning via Visual Knowledge Distillation

  • 利用中间视觉表征作为监督信号,分离并删除特定视觉知识。
  • 在保持文本能力的同时,显著提升删减效果与模型效率。
  • 首个评估多模态模型删减鲁棒性的方法,适合安全敏感场景。

近年来,机器删减方法被提出以消除训练好的大模型中的敏感信息。然而,现有方法多针对语言模型,面向多模态大模型(MLLM)的删减仍处于初级阶段。受近期对MLLM内部机制研究的启发,我们提出将嵌入在MLLM中的视觉与文本知识解耦,并设计专用方法选择性擦除目标视觉知识,同时保留文本知识。不同于依赖输出层监督的已有方法,本工作引入视觉知识蒸馏(VKD)机制,利用MLLM内部的中间视觉表示作为监督信号,显著提升删减效果与模型实用性。此外,由于仅微调视觉组件,该方法具有显著的计算效率优势。大量实验表明,本方法在有效性和效率上均优于当前最优删减方法。更重要的是,我们首次评估了MLLM删减对重学习攻击的鲁棒性。

原文摘要 · Abstract (English)

Recently, machine unlearning approaches have been proposed to remove sensitive information from well-trained large models. However, most existing methods are tailored for LLMs, while MLLM-oriented unlearning remains at its early stage. Inspired by recent studies exploring the internal mechanisms of MLLMs, we propose to disentangle the visual and textual knowledge embedded within MLLMs and introduce a dedicated approach to selectively erase target visual knowledge while preserving textual knowledge. Unlike previous unlearning methods that rely on output-level supervision, our approach introduces a Visual Knowledge Distillation (VKD) scheme, which leverages intermediate visual representations within the MLLM as supervision signals. This design substantially enhances both unlearning effectiveness and model utility. Moreover, since our method only fine-tunes the visual components of the MLLM, it offers significant efficiency advantages. Extensive experiments demonstrate that our approach outperforms state-of-the-art unlearning methods in terms of both effectiveness and efficiency. Moreover, we are the first to evaluate the robustness of MLLM unlearning against relearning attacks.

多模态模型知识蒸馏隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。