用强化学习精准删除文本生成图像中的危险内容,不伤及正常语义。
ForceForget: Reinforcement Concept Removal for Enhancing Safety in Text-to-Image Models

- 通过强化学习优化概念移除奖励,实现安全与可用性的平衡。
- 在多个数据集上显著减少有害生成,同时保持良性图像高保真度。
- 适用于新兴的图生图场景,可扩展至艺术风格等通用概念移除。
随着生成式AI的发展,文本到图像(T2I)模型具备生成多样化内容的能力,但依然可能生成不安全内容。现有概念擦除方法常过度删除危险概念,并抑制包含于有害提示中的良性概念,影响模型实用性。本文聚焦于在消除不安全内容的同时,维持模型对安全语义的准确理解,通过强化学习优化概念擦除奖励(CER)。为避免过度内容删除,引入安全适配器(Safe Adapter),对部分文本嵌入进行投影,实现跨注意力层的高效概念调控。在多个数据集上的大量实验表明,该方法在缓解不安全内容生成方面优于现有最先进(SOTA)概念擦除方法,同时保持良性图像的高保真度。在鲁棒性方面,本方法在对抗红队工具测试中表现更优。此外,该方法在新兴的图像到图像(I2I)场景中也更有效。最后,我们还将方法扩展至移除一般概念,如艺术风格和物体。免责声明:本文涉及可能令部分读者不适的色情内容讨论,所有图像均为合成或来自公开数据集。
原文摘要 · Abstract (English)
With the advance of generative AI, the text-to-image (T2I) model has the ability to generate various contents. However, T2I models still can generate unsafe contents. To alleviate this issue, various concept erasing methods are proposed. However, existing methods tend to excessively erase unsafe concepts and suppress benign concepts contained in harmful prompts, which can negatively affect model utility. In this paper, we focus on eliminating unsafe content while maintaining model capability in safe semantic meaning interpretation by optimizing the concept erasing reward (CER) with reinforcement learning. To avoid overly content erasure, we introduce the Safe Adapter to project partial text embedding for efficient concept regulation in cross-attention layers. Extensive experiments conducted on different datasets demonstrate the effectiveness of the proposed method in alleviating unsafe content generation while preserving the high fidelity of benign images compared with existing state-of-the-art (SOTA) concept erasing methods. In terms of robustness, our method outperforms counterparts against red-teaming tools. Moreover, we showcase the proposed approach is more effective in emerging image-to-image (I2I) scenarios compared with others. Lastly, we extend our method to erase general concepts, such as artistic styles and objects. Disclaimer: This paper includes discussions of sexually explicit content that may be offensive to certain readers. All images used in this work are synthesized or from public datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。