可批量擦除上千个概念,让AI图像生成更安全可控。
Erasing Thousands of Concepts: Towards Scalable and Practical Concept Erasure for Text-to-Image Diffusion Models
- 用t分布混合模型建模概念分布,精准定位目标内容。
- 支持2000多个概念擦除,保持图像生成质量不下降。
- 抗攻击能力强,适合实际部署在主流文生图模型中。
大规模文本到图像(T2I)扩散模型虽生成效果出色,但可能重现版权内容等有害信息。概念擦除成为缓解风险的策略,但现有方法难以兼顾可扩展性、精确性和鲁棒性,仅能处理数百个概念。为此,我们提出可擦除数千概念的框架ETC。首先通过学生t分布混合模型(tMM)建模低秩概念分布,利用仿射最优传输实现精准擦除,同时通过锚定目标概念边界而不依赖预设锚点,保护其他概念。随后训练基于专家混合(MoE)的MoEraser模块,移除目标嵌入并保留锚点嵌入。通过向文本嵌入投影器注入噪声,并微调MoEraser以恢复,使系统具备对白盒攻击(如模块移除)的鲁棒性。在跨异构领域和多种扩散模型上针对超过2000个概念的实验表明,ETC在大规模概念擦除中达到最先进水平,兼具可扩展性与高精度。
原文摘要 · Abstract (English)
Large-scale text-to-image (T2I) diffusion models deliver remarkable visual fidelity but pose safety risks due to their capacity to reproduce undesirable content, such as copyrighted ones. Concept erasure has emerged as a mitigation strategy, yet existing approaches struggle to balance scalability, precision, and robustness, which restricts their applicability to erasing only a few hundred concepts. To address these limitations, we present Erasing Thousands of Concepts (ETC), a scalable framework capable of erasing thousands of concepts while preserving generation quality. Our method first models low-rank concept distributions via a Student's t-distribution Mixture Model (tMM). It enables pin-point erasure of target concepts via affine optimal transport while preserving others by anchoring the boundaries of target concept distributions without pre-defined anchor concepts. We then train a Mixture-of-Experts (MoE)-based module, termed MoEraser, which removes target embeddings while preserving the anchor embeddings. By injecting noise into the text embedding projector and fine-tuning MoEraser for recovery, our framework achieves robustness to white-box attack such as module removal. Extensive experiments on over 2,000 concepts across heterogeneous domains and diffusion models demerate state-of-the-art scalability and precision in large-scale concept erasure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。