让视觉语言模型精准删除特定概念,不伤及无关信息。
ICED: Concept-level Machine Unlearning via Interpretable Concept Decomposition

- 用大模型构建任务专属概念词表,分解图像语义
- 通过稀疏非负组合实现概念级知识操控
- 可精准遗忘目标概念,同时保留其他语义和模型性能
视觉语言模型中的机器遗忘通常在图像或实例层面进行,难以精确删除目标知识而不影响无关语义。由于单张图像常包含多个纠缠的概念,包括需遗忘的目标概念与应保留的上下文信息,这一问题尤为突出。本文提出一种可解释的概念级遗忘框架,利用多模态大语言模型从遗忘数据集中构建紧凑的任务专属概念词汇表。除模态对齐外,视觉表示被分解为稀疏、非负的语义概念组合,提供细粒度知识操作的显式接口。基于此分解,方法将遗忘建模为概念级优化:选择性抑制目标概念,同时保留实例内非目标语义与跨模态全局知识。在域内与域外遗忘设置下的大量实验表明,该方法实现了更全面的目标遗忘,更好保留同一图像中的非目标知识,并在模型效用上优于现有VLM遗忘方法。
原文摘要 · Abstract (English)
Machine unlearning in Vision-Language Models (VLMs) is typically performed at the image or instance level, making it difficult to precisely remove target knowledge without affecting unrelated semantics. This issue is especially pronounced since a single image often contains multiple entangled concepts, including both target concepts to be forgotten and contextual information that should be preserved. In this paper, we propose an interpretable concept-level unlearning framework for VLMs, which constructs a compact task-specific concept vocabulary from the forgetting set using a multimodal large language model. In addition to modality alignment, visual representations are decomposed into sparse, nonnegative combinations of semantic concepts, providing an explicit interface for fine-grained knowledge manipulation. Based on this decomposition, our method formulates unlearning as concept-level optimization, where target concepts are selectively suppressed while intra-instance non-target semantics and global cross-modal knowledge are preserved. Extensive experiments across both in-domain and out-of-domain forgetting settings demonstrate that our method enables more comprehensive target forgetting, better preserves non-target knowledge within the same image, and maintains competitive model utility compared with existing VLM unlearning methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。