用大规模三元组数据训练,让视频风格迁移更稳定、一致。
VISTA: Triplet-Supervised Video Style Transfer with Diffusion Transformers
- 构建三元组数据集,联合建模风格、内容与运动
- 在1000种风格上实现最高风格保真度和时序一致性
- 适合需要高质量视频风格迁移的研究者与开发者
视频风格迁移旨在保持内容、结构和动作的前提下,将视频渲染为指定艺术风格。尽管图像风格化已取得显著进展,但视频风格化仍面临时序不一致的挑战。现有方法多对帧或关键帧进行风格化,并通过启发式时序传播保证一致性,但在遮挡、非遮挡及长期运动下易出现漂移与闪烁。我们指出根本瓶颈在于缺乏大规模三元组数据和联合建模风格、内容与运动的训练范式。为此,我们提出VISTA-1000,一个包含1000种风格的合成数据集,包含风格参考、原始视频与风格化视频的运动对齐三元组,并设计基于扩散变压器的上下文视频风格迁移框架,配备轻量级风格适配器以实现鲁棒风格提取。大量实验表明,该方法在风格保真度、时序一致性和内容保留方面达到当前最优水平。
原文摘要 · Abstract (English)
Video style transfer aims to render videos in a target artistic style while preserving content, structure, and motion. While image stylization has advanced rapidly, video stylization remains challenging due to temporal inconsistency. Most existing methods stylize frames or keyframes and enforce consistency via heuristic temporal propagation, which is brittle under occlusions, disocclusions, and long-term motion, leading to drift and flickering artifacts. We argue that a fundamental bottleneck lies in the lack of large-scale triplet data and a principled training paradigm that jointly models and disentangles style, content, and motion.To address this, we introduce VISTA-1000, a synthetic dataset with 1,000 styles and motion-aligned triplets of style reference, clean video, and stylized video, and propose a diffusion-transformer-based in-context video style transfer framework with a lightweight style adapter for robust style extraction. Extensive experiments demonstrate SOTA performance in style fidelity, temporal consistency, and content preservation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。