统一视频风格化框架,支持文本、图像、首帧三种风格控制。
DreamStyle: A Unified Framework for Video Stylization
- 基于I2V模型,用特定令牌矩阵减少风格混淆。
- 在三类风格控制下均实现高一致性与高质量视频生成。
- 自建数据集缓解风格不一致与时间闪烁问题,适合内容创作者。
视频风格化是视频生成模型的重要下游任务,但尚未充分探索。现有风格条件包括文本、风格图像和风格化首帧,各有优势:文本灵活,图像提供精确视觉锚点,首帧使长视频风格化成为可能。然而,现有方法多局限于单一风格条件,应用范围受限;且缺乏高质量数据集,导致风格不一致与时间闪烁。为此,我们提出DreamStyle,一种统一的视频风格化框架,支持文本引导、风格图像引导和首帧引导三种方式,并设计了高质量配对视频数据采集流程。DreamStyle基于基础图像到视频(I2V)模型,采用低秩适配(LoRA)结合令牌特异性上投影矩阵,有效降低不同条件令牌间的混淆。定性与定量评估表明,DreamStyle在三项任务中均表现优异,优于现有方法,在风格一致性和视频质量方面显著提升。
原文摘要 · Abstract (English)
Video stylization, an important downstream task of video generation models, has not yet been thoroughly explored. Its input style conditions typically include text, style image, and stylized first frame. Each condition has a characteristic advantage: text is more flexible, style image provides a more accurate visual anchor, and stylized first frame makes long-video stylization feasible. However, existing methods are largely confined to a single type of style condition, which limits their scope of application. Additionally, their lack of high-quality datasets leads to style inconsistency and temporal flicker. To address these limitations, we introduce DreamStyle, a unified framework for video stylization, supporting (1) text-guided, (2) style-image-guided, and (3) first-frame-guided video stylization, accompanied by a well-designed data curation pipeline to acquire high-quality paired video data. DreamStyle is built on a vanilla Image-to-Video (I2V) model and trained using a Low-Rank Adaptation (LoRA) with token-specific up matrices that reduces the confusion among different condition tokens. Both qualitative and quantitative evaluations demonstrate that DreamStyle is competent in all three video stylization tasks, and outperforms the competitors in style consistency and video quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。