通过频域替换实现高效可控的文本驱动图像迁移。
FBSDiff++: Improved Frequency Band Substitution of Diffusion Features for Efficient and Highly Controllable Text-Driven Image-to-Image Translation
- 从频域角度动态替换扩散特征,实现无训练迁移。
- 推理速度提升8.9倍,支持任意分辨率输入。
- 可精细控制风格、布局与轮廓,适合创意设计者。
随着大规模文本到图像(T2I)扩散模型在开放领域图像生成中取得显著进展,越来越多研究关注其向文本驱动图像到图像(I2I)迁移的自然延伸。在此任务中,源图像提供视觉引导,文本提示提供语义引导。本文提出FBSDiff,一种新颖框架,从频域视角将现成T2I扩散模型适配至I2I范式。通过动态替换扩散特征的频率带,实现无需训练、微调或在线优化的即插即用式多模态控制:逐步替换低频、中频、高频带,分别实现外观、布局、轮廓引导的I2I迁移。同时,通过调节替换频带宽度,可连续控制迁移强度。为进一步提升效率、灵活性和功能,本文提出FBSDiff++,主要改进包括:(1) 通过优化模型架构实现推理速度大幅提升(8.9×加速);(2) 改进频带替换模块,支持任意分辨率与长宽比的输入图像;(3) 扩展功能,仅需微调即可实现局部图像编辑与风格化内容生成。大量定性和定量实验验证了FBSDiff++在视觉质量、效率、多功能性与可控性方面优于现有先进方法。
原文摘要 · Abstract (English)
With large-scale text-to-image (T2I) diffusion models achieving significant advancements in open-domain image creation, increasing attention has been focused on their natural extension to the realm of text-driven image-to-image (I2I) translation, where a source image acts as visual guidance to the generated image in addition to the textual guidance provided by the text prompt. We propose FBSDiff, a novel framework adapting off-the-shelf T2I diffusion model into the I2I paradigm from a fresh frequency-domain perspective. Through dynamic frequency band substitution of diffusion features, FBSDiff realizes versatile and highly controllable text-driven I2I in a plug-and-play manner (without need for model training, fine-tuning, or online optimization), allowing appearance-guided, layout-guided, and contour-guided I2I translation by progressively substituting low-frequency band, mid-frequency band, and high-frequency band of latent diffusion features, respectively. In addition, FBSDiff flexibly enables continuous control over I2I correlation intensity simply by tuning the bandwidth of the substituted frequency band. To further promote image translation efficiency, flexibility, and functionality, we propose FBSDiff++ which improves upon FBSDiff mainly in three aspects: (1) accelerate inference speed by a large margin (8.9$\times$ speedup in inference) with refined model architecture; (2) improve the Frequency Band Substitution module to allow for input source images of arbitrary resolution and aspect ratio; (3) extend model functionality to enable localized image manipulation and style-specific content creation with only subtle adjustments to the core method. Extensive qualitative and quantitative experiments verify superiority of FBSDiff++ in I2I translation visual quality, efficiency, versatility, and controllability compared to related advanced approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。