无需微调,一张图就能生成指定风格和动作的视频。
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner
- 用合成数据集训练模型,实现零样本主体视频定制。
- 生成视频在身份保留、动态表现和文本对齐上均领先。
- 适合需要快速定制视频内容的创作者或设计师。
我们提出SUGAR,一种零样本的主体驱动视频定制方法。给定一张输入图像,SUGAR能够生成图像中主体的视频,并根据用户输入的文本指令对齐任意视觉属性(如风格、动作)。与以往需测试时微调或无法生成文本对齐视频的方法不同,SUGAR在不增加测试成本的情况下实现更优效果。为实现零样本能力,我们构建了一个可扩展的合成数据集,包含250万组图像-视频-文本三元组。此外,我们提出多种改进方法,包括特殊注意力设计、优化训练策略和精炼采样算法。大量实验表明,SUGAR在身份保留、视频动态性和视频-文本对齐方面均达到当前最优性能,验证了该方法的有效性。
原文摘要 · Abstract (English)
We present SUGAR, a zero-shot method for subject-driven video customization. Given an input image, SUGAR is capable of generating videos for the subject contained in the image and aligning the generation with arbitrary visual attributes such as style and motion specified by user-input text. Unlike previous methods, which require test-time fine-tuning or fail to generate text-aligned videos, SUGAR achieves superior results without the need for extra cost at test-time. To enable zero-shot capability, we introduce a scalable pipeline to construct synthetic dataset which is specifically designed for subject-driven customization, leading to 2.5 millions of image-video-text triplets. Additionally, we propose several methods to enhance our model, including special attention designs, improved training strategies, and a refined sampling algorithm. Extensive experiments are conducted. Compared to previous methods, SUGAR achieves state-of-the-art results in identity preservation, video dynamics, and video-text alignment for subject-driven video customization, demonstrating the effectiveness of our proposed method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。