arXiv:2603.00436cs.LGcs.AI2026-03

提出防御知识遗忘攻击的新方法,让模型删数据不伤其他知识。

ROKA: Robust Knowledge Unlearning against Adversaries

  • 将神经网络视为知识系统,通过神经修复实现有损保留的删减。
  • 在多种大模型上验证,删数据后性能不降反升,抵御间接攻击。
  • 首次提供知识保留的理论保障,适合隐私保护场景使用。

机器学习中的知识遗忘对数据隐私至关重要,但现有方法常导致知识污染,使模型性能下降,进而被用于新型推理和后门攻击。多数研究依赖数据投毒或复制来设计对抗性遗忘请求。本文提出一种新型间接遗忘攻击,无需操纵数据,而是利用知识污染后果干扰模型在安全关键任务上的准确率。为应对该威胁,我们建立神经网络作为神经知识系统的理论框架,提出ROKA策略,聚焦于神经修复。不同于仅删除信息的传统方法,ROKA通过消除遗忘数据的影响并强化概念邻近知识,实现建设性再平衡。据我们所知,这是首个在遗忘过程中提供知识保留理论保证的工作。在视觉变换器、多模态模型及大语言模型上的评估表明,ROKA能有效遗忘目标数据,同时保持甚至提升留存数据的准确性,从而有效缓解间接遗忘攻击。

原文摘要 · Abstract (English)

The need for machine unlearning is critical for data privacy, yet existing methods often cause Knowledge Contamination by unintentionally damaging related knowledge. Such a degraded model performance after unlearning has been recently leveraged for new inference and backdoor attacks. Most studies design adversarial unlearning requests that require poisoning or duplicating training data. In this study, we introduce a new unlearning-induced attack model, namely indirect unlearning attack, which does not require data manipulation but exploits the consequence of knowledge contamination to perturb the model accuracy on security-critical predictions. To mitigate this attack, we introduce a theoretical framework that models neural networks as Neural Knowledge Systems. Based on this, we propose ROKA, a robust unlearning strategy centered on Neural Healing. Unlike conventional unlearning methods that only destroy information, ROKA constructively rebalances the model by nullifying the influence of forgotten data while strengthening its conceptual neighbors. To the best of our knowledge, our work is the first to provide a theoretical guarantee for knowledge preservation during unlearning. Evaluations on various large models, including vision transformers, multi-modal models, and large language models, show that ROKA effectively unlearns targets while preserving, or even enhancing, the accuracy of retained data, thereby mitigating the indirect unlearning attacks.

知识遗忘模型安全神经修复

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。