无需训练,多参考帧实现高质量视频风格化
FreeViS: Training-free Video Stylization with Inconsistent References
- 融合多个风格参考,提升风格细节与时间一致性
- 避免闪烁抖动,保持低显著区域的纹理清晰
- 适合快速生成高质视频风格,无需数据与计算开销
视频风格化在内容创作中至关重要,但存在挑战。直接逐帧应用图像风格化会破坏时间连贯性并降低风格丰富度。现有训练方法需成对视频数据且计算成本高。本文提出FreeViS,一种无需训练的视频风格化框架,通过将多个风格参考整合至预训练图像到视频(I2V)模型,有效缓解传播误差,避免闪烁和卡顿。同时结合高频补偿约束内容布局与运动,并利用基于光流的运动线索保留低显著区域的风格纹理。大量实验表明,FreeViS在风格保真度和时间一致性上均优于近期基线,获得更强人类偏好。其无需训练的流程为高质量、时间一致的视频风格化提供了高效经济方案。代码与示例视频见https://xujiacong.github.io/FreeViS/
原文摘要 · Abstract (English)
Video stylization plays a key role in content creation, but it remains a challenging problem. Naïvely applying image stylization frame-by-frame hurts temporal consistency and reduces style richness. Alternatively, training a dedicated video stylization model typically requires paired video data and is computationally expensive. In this paper, we propose FreeViS, a training-free video stylization framework that generates stylized videos with rich style details and strong temporal coherence. Our method integrates multiple stylized references to a pretrained image-to-video (I2V) model, effectively mitigating the propagation errors observed in prior works, without introducing flickers and stutters. In addition, it leverages high-frequency compensation to constrain the content layout and motion, together with flow-based motion cues to preserve style textures in low-saliency regions. Through extensive evaluations, FreeViS delivers higher stylization fidelity and superior temporal consistency, outperforming recent baselines and achieving strong human preference. Our training-free pipeline offers a practical and economic solution for high-quality, temporally coherent video stylization. The code and videos can be accessed via https://xujiacong.github.io/FreeViS/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。