arXiv:2605.11592cs.LGcs.AI2026-05

首次系统分析模型去记忆技术,揭示其防御漏洞与理论边界。

SoK: Unlearnability and Unlearning for Model Dememorization

论文配图:SoK: Unlearnability and Unlearning for Model Dememorization
图 1 · 摘自论文原文
  • 构建统一分类框架,整合不可学习性与模型遗忘方法
  • 实验证明现有方法易受浅层遗忘干扰,存在虚假安全
  • 提出首个认证遗忘的深度理论保证,适合隐私保护研究者

先进的模型去记忆方法,包括可用性投毒(不可学习性)和机器遗忘,正成为应对机器学习中数据滥用的关键防护手段。训练阶段的不可学习性在数据发布前嵌入难以察觉的扰动以降低可学习性;训练后的遗忘则移除模型中已习得的信息,防止未经授权的披露或使用。尽管两者均旨在维护知识隐匿权,但其脆弱性与共通基础尚不明确。具体而言,不可学习性与遗忘均面临浅层去记忆问题,导致虚假的数据可学习性降低或权重扰动下的误遗忘。此外,输入扰动可能影响下游遗忘效果,而遗忘过程可能意外恢复被不可学习性隐藏的领域知识。这种相互作用亟需深入探究。最后,现有防御缺乏形式化保证,无法提供对浅层去记忆的理论洞察。本文首次系统化分析了基于不可学习性与遗忘的模型去记忆方法。贡献有三:(i) 构建不可学习性与可扩展遗忘方法的统一分类体系;(ii) 实证评估揭示主流方法的鲁棒性、相互作用与浅层去记忆缺陷;(iii) 首次为经认证遗忘处理的模型提供去记忆深度的理论保证。这些成果为贯穿机器学习全生命周期的去记忆机制融合奠定基础,实现敏感知识更深层次的不可记忆状态。

原文摘要 · Abstract (English)

Advanced model dememorization methods, including availability poisoning (unlearnability) and machine unlearning, are emerging as key safeguards against data misuse in machine learning (ML). At the training stage, unlearnability embeds imperceptible perturbations into data before release to reduce learnability. At the post-training stage, unlearning removes previously acquired information from models to prevent unauthorized disclosure or use. While both defenses aim to preserve the right to withhold knowledge, their vulnerabilities and shared foundations remain unclear. Specifically, both unlearnability and unlearning suffer from issues such as shallow dememorization, leading to falsely claimed data learnability reduction or forgetting in the presence of weight perturbations. Moreover, input perturbations may affect the effectiveness of downstream unlearning, while unlearning may inadvertently recover domain knowledge hidden by unlearnability. This interplay calls for deeper investigation. Finally, there is a lack of formal guarantees to provide theoretical insights into current defenses against shallow dememorization. In this Systematization of Knowledge, we present the first integrated analysis of model dememorization approaches leveraging unlearnability and unlearning. Our contributions are threefold: (i) a unified taxonomy of unlearnability and scalable unlearning methods; (ii) an empirical evaluation revealing the robustness, interplay, and shallow dememorization of leading methods; and (iii) the first theoretical guarantee on dememorization depth for models processed through certified unlearning. These results lay the foundation for unifying dememorization mechanisms across the ML lifecycle to achieve a deeper immemor state for sensitive knowledge.

模型遗忘隐私保护不可学习性理论保证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。