arXiv:2602.06331cs.LG2026-02

让模型遗忘特定类别时,仍能准确识别异常数据。

Don't Break the Boundary: Continual Unlearning for OOD Detection Based on Free Energy Repulsion

  • 用能量排斥机制把被遗忘类推到异常区域,保留正常数据结构
  • 遗忘后对剩余类和外部异常数据的检测能力几乎不下降
  • 适合需要持续删除数据隐私的AI系统,如医疗或金融应用

在开放世界中部署可信AI面临双重挑战:需具备强大的分布外(OOD)检测能力以保障安全,同时要求灵活的机器遗忘功能以满足隐私合规与模型修正需求。然而,这一目标存在根本性几何矛盾:现有OOD检测依赖于静态紧凑的数据流形,而传统以分类为导向的遗忘方法会破坏此结构,导致模型在移除目标类别时灾难性丧失异常检测能力。为此,我们首次定义边界保持型类遗忘问题,并提出关键概念转变:在OOD检测背景下,有效遗忘等价于将目标类别转化为分布外样本。基于此,提出总自由能排斥(TFER)框架。受自由能原理启发,TFER构建新颖的‘推-拉’博弈机制:通过拉力将保留类别锚定在低能量内域流形中,同时利用自由能排斥力主动将遗忘类别驱逐至高能量分布外区域。该方法通过参数高效微调实现,避免全量重训的高昂成本。大量实验表明,TFER在精确遗忘的同时,最大程度保留了模型对剩余类别及外部OOD数据的判别性能。更重要的是,研究揭示TFER特有的推拉平衡赋予模型内在结构稳定性,无需额外约束即可有效抵抗灾难性遗忘,展现出在持续遗忘任务中的卓越潜力。

原文摘要 · Abstract (English)

Deploying trustworthy AI in open-world environments faces a dual challenge: the necessity for robust Out-of-Distribution (OOD) detection to ensure system safety, and the demand for flexible machine unlearning to satisfy privacy compliance and model rectification. However, this objective encounters a fundamental geometric contradiction: current OOD detectors rely on a static and compact data manifold, whereas traditional classification-oriented unlearning methods disrupt this delicate structure, leading to a catastrophic loss of the model's capability to discriminate anomalies while erasing target classes. To resolve this dilemma, we first define the problem of boundary-preserving class unlearning and propose a pivotal conceptual shift: in the context of OOD detection, effective unlearning is mathematically equivalent to transforming the target class into OOD samples. Based on this, we propose the TFER (Total Free Energy Repulsion) framework. Inspired by the free energy principle, TFER constructs a novel Push-Pull game mechanism: it anchors retained classes within a low-energy ID manifold through a pull mechanism, while actively expelling forgotten classes to high-energy OOD regions using a free energy repulsion force. This approach is implemented via parameter-efficient fine-tuning, circumventing the prohibitive cost of full retraining. Extensive experiments demonstrate that TFER achieves precise unlearning while maximally preserving the model's discriminative performance on remaining classes and external OOD data. More importantly, our study reveals that the unique Push-Pull equilibrium of TFER endows the model with inherent structural stability, allowing it to effectively resist catastrophic forgetting without complex additional constraints, thereby demonstrating exceptional potential in continual unlearning tasks.

模型遗忘异常检测自由能持续学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。