无需微调即可实现高质量跨语义图像渐变,速度超快。
FreeMorph: Tuning-Free Generalized Image Morphing with Diffusion Model
- 通过改进自注意力机制,动态融合双输入引导信号
- 生成序列保持身份一致且过渡方向可控,速度比现有方法快10到50倍
- 适合需要快速、通用图像变形的创意设计与视频制作场景
我们提出FreeMorph,首个无需微调的通用图像渐变方法,可处理语义或布局不同的输入。现有方法依赖预训练扩散模型微调,受限于时间及语义/布局差异。FreeMorph在不进行实例训练的前提下实现高保真渐变。尽管无微调方法因多步去噪过程的非线性及预训练模型固有偏差面临质量挑战,本文通过两项创新解决:1)提出引导感知球面插值,通过修改自注意力模块显式引入输入图像引导,缓解身份丢失并确保生成序列方向一致;2)设计步进导向变化趋势,融合来自每张输入图像的自注意力模块,实现受控且一致的过渡。大量实验表明,FreeMorph优于现有方法,速度提升10至50倍,建立图像渐变新基准。
原文摘要 · Abstract (English)
We present FreeMorph, the first tuning-free method for image morphing that accommodates inputs with different semantics or layouts. Unlike existing methods that rely on finetuning pre-trained diffusion models and are limited by time constraints and semantic/layout discrepancies, FreeMorph delivers high-fidelity image morphing without requiring per-instance training. Despite their efficiency and potential, tuning-free methods face challenges in maintaining high-quality results due to the non-linear nature of the multi-step denoising process and biases inherited from the pre-trained diffusion model. In this paper, we introduce FreeMorph to address these challenges by integrating two key innovations. 1) We first propose a guidance-aware spherical interpolation design that incorporates explicit guidance from the input images by modifying the self-attention modules, thereby addressing identity loss and ensuring directional transitions throughout the generated sequence. 2) We further introduce a step-oriented variation trend that blends self-attention modules derived from each input image to achieve controlled and consistent transitions that respect both inputs. Our extensive evaluations demonstrate that FreeMorph outperforms existing methods, being 10x ~ 50x faster and establishing a new state-of-the-art for image morphing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。