arXiv:2602.17088cs.LG2026-02

通过语义引导重对齐,实现精准删去特定数据影响而不损害模型性能。

MeGU: Machine-Guided Unlearning with Target Feature Disentanglement

  • 用多模态大模型生成语义扰动标签,指导目标样本的特征重对齐。
  • 引入正负噪声对分离目标概念特征,保留共享语义结构。
  • 适合需要精准删除用户数据且保护模型可用性的场景。

训练数据隐私问题日益突出,“被遗忘的权利”成为关键需求,推动有效机器删忆的发展。然而现有方法普遍存在根本性权衡:激进删忆会损害模型在保留数据上的性能,保守策略则残留目标信息。本文分析预训练模型中的内在表征特性,发现语义类别概念在特征层面存在纠缠,共享相关特征但保留特定判别成分,这限制了传统删忆范式的效果。为此提出机器引导删忆(MeGU)框架,通过概念感知重对齐实现删忆。具体地,利用多模态大语言模型(MLLM)为目标样本分配语义有意义的扰动标签,确定重对齐方向;通过估计类间概念相似性构建轻量过渡矩阵提升效率;引入正负特征噪声对,显式解耦目标概念影响。微调中,负噪声抑制目标特异性特征模式,正噪声强化关联特征并对其与扰动概念对齐。该协同设计实现目标特征的选择性破坏,同时保留共享语义结构,从而支持可控、选择性遗忘,有效缓解删忆不足与过度删忆问题。

原文摘要 · Abstract (English)

The growing concern over training data privacy has elevated the "Right to be Forgotten" into a critical requirement, thereby raising the demand for effective Machine Unlearning. However, existing unlearning approaches commonly suffer from a fundamental trade-off: aggressively erasing the influence of target data often degrades model utility on retained data, while conservative strategies leave residual target information intact. In this work, the intrinsic representation properties learned during model pretraining are analyzed. It is demonstrated that semantic class concepts are entangled at the feature-pattern level, sharing associated features while preserving concept-specific discriminative components. This entanglement fundamentally limits the effectiveness of existing unlearning paradigms. Motivated by this insight, we propose Machine-Guided Unlearning (MeGU), a novel framework that guides unlearning through concept-aware re-alignment. Specifically, Multi-modal Large Language Models (MLLMs) are leveraged to explicitly determine re-alignment directions for target samples by assigning semantically meaningful perturbing labels. To improve efficiency, inter-class conceptual similarities estimated by the MLLM are encoded into a lightweight transition matrix. Furthermore, MeGU introduces a positive-negative feature noise pair to explicitly disentangle target concept influence. During finetuning, the negative noise suppresses target-specific feature patterns, while the positive noise reinforces remaining associated features and aligns them with perturbing concepts. This coordinated design enables selective disruption of target-specific representations while preserving shared semantic structures. As a result, MeGU enables controlled and selective forgetting, effectively mitigating both under-unlearning and over-unlearning.

机器删忆特征解耦大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。