arXiv:2410.15007cs.CVcs.MM2024-10中稿 · ACMMM Asia 2024被引 14

无需训练,用扩散模型实现内容与风格的精准平衡融合。

DiffuseST: Unleashing the Capability of the Diffusion Model for Style Transfer

  • 结合文本嵌入与空间特征,分离注入内容和风格信息。
  • 通过扩散模型逐步注入,提升内容保留与风格融合的平衡性。
  • 无需训练,可扩展至其他图像生成任务,适合追求可控生成的用户。

风格迁移旨在将风格图像的艺术表现与内容图像的结构信息融合。现有方法依赖特定网络或预训练模型学习内容与风格特征,但仅使用文本或空间表示,难以平衡内容与风格。本文提出一种全新的、无需训练的风格迁移方法,将文本嵌入与空间特征相结合,并分离内容与风格的注入过程。具体地,采用BLIP-2编码器提取风格图像的文本表示,利用DDIM反演技术获取内容与风格分支的中间嵌入作为空间特征。最终,借助扩散模型的逐步生成特性,在目标分支中分离注入内容与风格信息,从而提升内容保持与风格融合之间的平衡。大量实验表明,DiffuseST能实现均衡且可控的风格迁移效果,具备向其他任务扩展的潜力。

原文摘要 · Abstract (English)

Style transfer aims to fuse the artistic representation of a style image with the structural information of a content image. Existing methods train specific networks or utilize pre-trained models to learn content and style features. However, they rely solely on textual or spatial representations that are inadequate to achieve the balance between content and style. In this work, we propose a novel and training-free approach for style transfer, combining textual embedding with spatial features and separating the injection of content or style. Specifically, we adopt the BLIP-2 encoder to extract the textual representation of the style image. We utilize the DDIM inversion technique to extract intermediate embeddings in content and style branches as spatial features. Finally, we harness the step-by-step property of diffusion models by separating the injection of content and style in the target branch, which improves the balance between content preservation and style fusion. Various experiments have demonstrated the effectiveness and robustness of our proposed DiffeseST for achieving balanced and controllable style transfer results, as well as the potential to extend to other tasks.

风格迁移扩散模型无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。