无需微调,用自适应插值实现视频风格无缝过渡。
SOYO: A Tuning-Free Approach for Video Style Morphing via Style-Adaptive Interpolation in Diffusion Models
- 不微调预训练模型,通过注意力注入与AdaIN保持结构一致。
- 自适应采样调度使多风格间过渡更均匀,避免突变。
- 适合需要自然风格渐变的视频生成任务,如艺术创作。
扩散模型在图像和视频风格化方面取得了显著进展。然而,现有方法大多聚焦单风格迁移,而涉及多风格的视频风格化需实现帧间平滑过渡。我们将这种视频帧间的连续风格变化称为视频风格形态转换(video style morphing)。当前方法在处理此类过渡时,常导致结构不连贯和风格突变。为此,我们提出SOYO,一种基于扩散模型的新型视频风格形态转换框架。该方法不微调预训练文本到图像扩散模型,结合注意力注入与AdaIN,以保持结构一致性并实现帧间平滑风格转换。此外,我们发现直接使用线性等距插值会导致风格转换不平衡。为此,提出一种新颖的自适应采样调度器,用于两风格图像之间的动态调节。大量实验表明,SOYO在开放域视频风格形态转换中优于现有方法,更好保持了视频帧的结构连贯性,同时实现稳定平滑的风格过渡。
原文摘要 · Abstract (English)
Diffusion models have achieved remarkable progress in image and video stylization. However, most existing methods focus on single-style transfer, while video stylization involving multiple styles necessitates seamless transitions between them. We refer to this smooth style transition between video frames as video style morphing. Current approaches often generate stylized video frames with discontinuous structures and abrupt style changes when handling such transitions. To address these limitations, we introduce SOYO, a novel diffusion-based framework for video style morphing. Our method employs a pre-trained text-to-image diffusion model without fine-tuning, combining attention injection and AdaIN to preserve structural consistency and enable smooth style transitions across video frames. Moreover, we notice that applying linear equidistant interpolation directly induces imbalanced style morphing. To harmonize across video frames, we propose a novel adaptive sampling scheduler operating between two style images. Extensive experiments demonstrate that SOYO outperforms existing methods in open-domain video style morphing, better preserving the structural coherence of video frames while achieving stable and smooth style transitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。