无需训练,基于语义对齐实现图像视频风格迁移
Semantix: An Energy Guided Sampler for Semantic Style Transfer
- 利用预训练扩散模型的语义理解能力,设计能量引导采样器
- 在图像和视频上均超越现有最佳方法,实现跨媒体通用迁移
- 适合需要高质量跨媒体风格迁移的开发者与研究者
近期风格与外观迁移进展显著,但多数方法将全局风格与局部外观分离处理,忽视语义对应关系。同时,图像与视频任务常被孤立处理,缺乏整合。为此,我们提出新任务:语义风格迁移,即根据语义对应关系,将参考图像的风格与外观特征迁移到目标视觉内容。随后,我们提出无需训练的方法 Semantix——一种基于能量引导的采样器,可同时利用预训练扩散模型的语义理解能力,指导风格与外观迁移。作为采样器,Semantix 可无缝应用于图像与视频模型,实现跨媒体通用迁移。具体而言,通过 SDE 将参考图与上下文图/视频反向映射至噪声空间后,利用精心设计的能量函数引导采样过程,包含三个关键组件:风格特征引导、空间特征引导及语义距离正则项。实验表明,Semantix 在图像与视频领域均有效完成语义风格迁移,并优于现有最先进方法。
原文摘要 · Abstract (English)
Recent advances in style and appearance transfer are impressive, but most methods isolate global style and local appearance transfer, neglecting semantic correspondence. Additionally, image and video tasks are typically handled in isolation, with little focus on integrating them for video transfer. To address these limitations, we introduce a novel task, Semantic Style Transfer, which involves transferring style and appearance features from a reference image to a target visual content based on semantic correspondence. We subsequently propose a training-free method, Semantix an energy-guided sampler designed for Semantic Style Transfer that simultaneously guides both style and appearance transfer based on semantic understanding capacity of pre-trained diffusion models. Additionally, as a sampler, Semantix be seamlessly applied to both image and video models, enabling semantic style transfer to be generic across various visual media. Specifically, once inverting both reference and context images or videos to noise space by SDEs, Semantix utilizes a meticulously crafted energy function to guide the sampling process, including three key components: Style Feature Guidance, Spatial Feature Guidance and Semantic Distance as a regularisation term. Experimental results demonstrate that Semantix not only effectively accomplishes the task of semantic style transfer across images and videos, but also surpasses existing state-of-the-art solutions in both fields. The project website is available at https://huiang-he.github.io/semantix/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。