无需反演的视频风格迁移,提升内容一致性与效率
Inversion-Free Video Style Transfer with Trajectory Reset Attention Control and Content-Style Bridging
- 通过轨迹重置注意力控制,避免内容泄露
- 支持从精准保留内容到生动表达风格的多样输出
- 无需调参,适合高效部署于图像与视频场景
视频风格迁移旨在改变视频的视觉风格同时保持内容不变。以往方法在使用图像驱动方式时,常因内容泄露和风格错位而表现不佳。本文提出轨迹重置注意力控制(TRAC),通过重置去噪轨迹并施加注意力约束,显著提升内容一致性,并大幅降低计算开销。同时引入风格媒介(Style Medium)概念,弥合内容与风格之间的差距,实现更精确、和谐的风格传递。基于此,我们构建了一个无需调参的框架,为图像与视频风格迁移提供稳定、灵活且高效的解决方案。实验表明,该框架可生成从高度保真内容到极具表现力的鲜明风格的多样化结果。
原文摘要 · Abstract (English)
Video style transfer aims to alter the style of a video while preserving its content. Previous methods often struggle with content leakage and style misalignment, particularly when using image-driven approaches that aim to transfer precise styles. In this work, we introduce Trajectory Reset Attention Control (TRAC), a novel method that allows for high-quality style transfer while preserving content integrity. TRAC operates by resetting the denoising trajectory and enforcing attention control, thus enhancing content consistency while significantly reducing the computational costs against inversion-based methods. Additionally, a concept termed Style Medium is introduced to bridge the gap between content and style, enabling a more precise and harmonious transfer of stylistic elements. Building upon these concepts, we present a tuning-free framework that offers a stable, flexible, and efficient solution for both image and video style transfer. Experimental results demonstrate that our proposed framework accommodates a wide range of stylized outputs, from precise content preservation to the production of visually striking results with vibrant and expressive styles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。