通过向量修改文本嵌入,抑制图像生成中与关键词强绑定的不想要内容。
Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models

- 在文本嵌入空间引入增量向量,直接削弱特定词的强关联内容
- 零样本即可获取增量向量,且在个性化模型中实现更精准抑制
- 新方法显著优于现有技术,尤其在控制不想要元素方面
文本到图像(T2I)扩散模型在从文本提示生成高质量图像方面取得了显著进展。然而,这些模型仍难以抑制与特定词汇强纠缠的内容。例如,在生成“查理·卓别林”时,即使明确要求不包含“胡子”,“胡子”仍会持续出现,因为“胡子”与“查理·卓别林”概念高度耦合。为此,我们提出一种新方法,直接在扩散模型的文本嵌入空间中抑制此类纠缠内容。该方法引入一个增量向量,用于修改文本嵌入以弱化生成图像中不希望出现的内容,并进一步证明该增量向量可通过零样本方式轻松获得。此外,我们提出了选择性抑制增量向量(SSDV)方法,将增量向量融入交叉注意力机制,从而在潜在生成区域更有效地抑制不需要的内容。同时,通过优化增量向量,我们在个性化T2I模型中实现了比以往基线更精确的抑制。大量实验结果表明,该方法在定量和定性指标上均显著优于现有方法。
原文摘要 · Abstract (English)
Text-to-Image (T2I) diffusion models have made significant progress in generating diverse high-quality images from textual prompts. However, these models still face challenges in suppressing content that is strongly entangled with specific words. For example, when generating an image of "Charlie Chaplin", a "mustache" consistently appears even if explicitly instructed not to include it, as the concept of "mustache" is strongly entangled with "Charlie Chaplin". To address this issue, we propose a novel approach to directly suppress such entangled content within the text embedding space of diffusion models. Our method introduces a delta vector that modifies the text embedding to weaken the influence of undesired content in the generated image, and we further demonstrate that this delta vector can be easily obtained through a zero-shot approach. Furthermore, we propose a Selective Suppression with Delta Vector (SSDV) method to adapt delta vector into the cross-attention mechanism, enabling more effective suppression of unwanted content in regions where it would otherwise be generated. Additionally, we enabled more precise suppression in personalized T2I models by optimizing delta vector, which previous baselines were unable to achieve. Extensive experimental results demonstrate that our approach significantly outperforms existing methods, both in terms of quantitative and qualitative metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。