arXiv:2502.07466cs.CV2025-02

通过遮蔽图像特征中的元素,有效防止风格迁移中的内容泄露。

Less is More: Masking Elements in Image Condition Features Avoids Content Leakages in Style Transfer Diffusion Models

  • 遮蔽风格参考图的特定特征元素,实现内容与风格解耦。
  • 在多种风格下提升生成效果,避免无关内容流入生成结果。
  • 无需调整模型参数,适合快速部署于现有扩散模型中。

给定一张风格参考图像作为额外图像条件,文本到图像扩散模型能够生成兼具提示词内容与参考图像视觉风格的图像。然而,当前最先进方法常难以分离参考图像中的内容与风格,导致内容泄露等问题。为此,我们提出一种基于遮蔽的方法,无需调整任何模型参数即可高效解耦内容与风格。通过简单遮蔽风格参考图像特征中的特定元素,我们发现:适当减少条件输入(如丢弃若干图像特征元素)能有效阻止无关内容进入扩散模型,从而提升文本到图像扩散模型的风格迁移性能。本文从理论与实验两方面验证该发现。在多种风格上的大量实验表明,该遮蔽方法有效且支持理论结论。

原文摘要 · Abstract (English)

Given a style-reference image as the additional image condition, text-to-image diffusion models have demonstrated impressive capabilities in generating images that possess the content of text prompts while adopting the visual style of the reference image. However, current state-of-the-art methods often struggle to disentangle content and style from style-reference images, leading to issues such as content leakages. To address this issue, we propose a masking-based method that efficiently decouples content from style without the need of tuning any model parameters. By simply masking specific elements in the style reference's image features, we uncover a critical yet under-explored principle: guiding with appropriately-selected fewer conditions (e.g., dropping several image feature elements) can efficiently avoid unwanted content flowing into the diffusion models, enhancing the style transfer performances of text-to-image diffusion models. In this paper, we validate this finding both theoretically and experimentally. Extensive experiments across various styles demonstrate the effectiveness of our masking-based method and support our theoretical results.

风格迁移扩散模型特征解耦

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。