现有概念替换技术在图像生成中无法真正消除不当内容,需兼顾效果与保真度。
Do Concept Replacement Techniques Really Erase Unacceptable Concepts?

- 用图像到图像模型验证主流概念替换技术失效
- 发现替换后仍会生成不当内容,且破坏原有概念
- 提出AntiMirror方法,兼顾删除无效内容与保留原意
生成式模型(尤其是基于扩散的文本到图像模型)取得了显著成功,但避免生成包含不当概念(如攻击性内容、版权内容或名人肖像)仍具挑战。概念替换技术(CRTs)旨在通过“擦除”模型中的不当概念来应对该问题。近期,模型提供商开始提供以图像和文本提示为输入的图像编辑服务,即图像到图像(I2I)模型。本文首次利用I2I模型实证表明,当前最先进的CRTs实际上并未真正擦除不当概念。尽管这些技术在文本到图像(T2I)流程中有效,但在新兴的I2I场景中可能失效,凸显T2I与I2I设置间的差异。我们进一步指出,理想的CRT应替换不当概念的同时保留输入中其他概念,称之为保真度。现有研究忽视了这一关键要求。为此,我们提出采用定向图像编辑技术实现有效性和保真度并重,并引入新方法AntiMirror,证明其可行性。
原文摘要 · Abstract (English)
Generative models, particularly diffusion-based text-to-image (T2I) models, have demonstrated astounding success. However, aligning them to avoid generating content with unacceptable concepts (e.g., offensive or copyrighted content, or celebrity likenesses) remains a significant challenge. Concept replacement techniques (CRTs) aim to address this challenge, often by trying to "erase" unacceptable concepts from models. Recently, model providers have started offering image editing services which accept an image and a text prompt as input, to produce an image altered as specified by the prompt. These are known as image-to-image (I2I) models. In this paper, we first use an I2I model to empirically demonstrate that today's state-of-the-art CRTs do not in fact erase unacceptable concepts. Existing CRTs are thus likely to be ineffective in emerging I2I scenarios, despite their proven ability to remove unwanted concepts in T2I pipelines, highlighting the need to understand this discrepancy between T2I and I2I settings. Next, we argue that a good CRT, while replacing unacceptable concepts, should preserve other concepts specified in the inputs to generative models. We call this fidelity. Prior work on CRTs have neglected fidelity in the case of unacceptable concepts. Finally, we propose the use of targeted image-editing techniques to achieve both effectiveness and fidelity. We present such a technique, AntiMirror, and demonstrate its viability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。