arXiv:2603.00992cs.LG2026-03

提出无需补偿的文本到图像模型去学习方法,精准消除敏感知识

Compensation-free Machine Unlearning in Text-to-Image Diffusion Models by Eliminating the Mutual Information

  • 通过最小化互信息精准移除特定概念知识
  • 实现敏感内容清除同时保持其他生成质量无下降
  • 首次做到不依赖任何补偿机制的去学习

扩散模型强大的生成能力引发了对生成敏感或不当内容的隐私与安全担忧。为此,机器去学习(MU)——在扩散模型中常称为概念擦除(CE)——被提出以从模型参数中移除特定知识,同时保留无关知识。尽管已有进展,现有方法往往过度且无差别地删除,导致无辜生成质量显著下降。为保留模型效用,先前工作依赖补偿机制,即重新引入部分剩余数据或显式约束剩余概念与预训练模型之间的差异。然而我们发现,超出补偿范围的生成仍受影响,表明此类事后补偿对大规模生成模型的通用效用而言本质不足。因此,本文主张开发无需补偿的概念擦除方法,精确识别并消除不良知识,使对其他生成的影响最小化。技术上,我们提出MiM-MU,通过精心设计的互信息最小化实现高效计算,并保持其他概念的采样分布。大量实验表明,该方法在有效移除目标概念的同时,维持了其他概念的高质量生成,且首次实现了无需任何事后补偿。

原文摘要 · Abstract (English)

The powerful generative capabilities of diffusion models have raised growing privacy and safety concerns regarding generating sensitive or undesired content. In response, machine unlearning (MU) -- commonly referred to as concept erasure (CE) in diffusion models -- has been introduced to remove specific knowledge from model parameters meanwhile preserving innocent knowledge. Despite recent advancements, existing unlearning methods often suffer from excessive and indiscriminate removal, which leads to substantial degradation in the quality of innocent generations. To preserve model utility, prior works rely on compensation, i.e., re-assimilating a subset of the remaining data or explicitly constraining the divergence from the pre-trained model on remaining concepts. However, we reveal that generations beyond the compensation scope still suffer, suggesting such post-remedial compensations are inherently insufficient for preserving the general utility of large-scale generative models. Therefore, in this paper, we advocate for developing compensation-free concept erasure operations, which precisely identify and eliminate the undesired knowledge such that the impact on other generations is minimal. In technique, we propose to MiM-MU, which is to unlearn a concept by minimizing the mutual information with a delicate design for computational effectiveness and for maintaining sampling distribution for other concepts. Extensive evaluations demonstrate that our proposed method achieves effective concept removal meanwhile maintaining high-quality generations for other concepts, and remarkably, without relying on any post-remedial compensation for the first time.

去学习扩散模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。