arXiv:2606.25465cs.CVcs.AI2026-06

通过逆向合成构建大规模数据集,实现长视频高质量文本驱动风格化。

EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis

论文配图:EchoStyle: Unlocking High-Fidelity Video Stylization with Reverse Data Synthesis
图 1 · 摘自论文原文
  • 采用视频到视频架构融合内容与文本风格,避免内容泄露。
  • 构建20,000对高质量视频风格化数据集,解决数据稀缺问题。
  • 支持任意长度视频,适合艺术创作与影视后期应用。

尽管图像风格化已得到广泛研究,视频风格化仍是智能内容创作中的关键难题。现有方法通常以参考图像作为风格先验,存在内容泄露、数据稀缺及对长视频适应性差等问题,导致严重风格漂移和运动失真。为此,我们提出EchoStyle,一种可扩展的文本驱动框架,实现任意长度视频的高质量风格化。首先,构建视频到视频架构,合理融合视频内容与文本风格。为解决数据稀缺问题,首创自动逆向合成流水线,建立包含20,000对高质量视频对的V-Style20k数据集。为支持长视频风格化,设计初始化-跟随模式与滑动窗口推理策略。大量实验表明,EchoStyle在多种艺术风格下表现优异,甚至可媲美领先闭源方案。

原文摘要 · Abstract (English)

While image stylization has been studied extensively, video stylization remains a critical and largely unsolved challenge in the field of intelligent content creation. Existing methods, usually utilizing a reference image as the style prior, suffer from content leakage, data scarcity and limited adaptability to long videos, leading to suboptimal results with severe style drift and motion distortion. For these issues, we present EchoStyle, a scalable text-driven framework to achieve high-quality stylization of videos with arbitrary lengths. To start with, we construct a video-to-video architecture to appropriately re-fuse the video content and the text style. To address data scarcity, we pioneer an automatic reverse-synthesis pipeline to establish V-Style20k, a large-scale stylization dataset of 20k high-quality video pairs. To facilitate long video stylization, we devise an init-follow-mode mechanism along with a sliding-window inference strategy. Extensive experiments demonstrate EchoStyle's excellent performance across a wide range of artistic styles, even comparable to leading closed-source solutions.

视频风格化文本驱动数据合成长视频处理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。