无需训练即可灵活用图片控制视频生成,支持任意数量和位置的图像条件。
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
- 通过反演图像到隐空间,利用随机块替换策略注入视觉特征。
- 动态调节条件强度,平衡视频创意与还原度,性能显著优于现有方法。
- 适用于UNet和Transformer架构,可推广至多种基础视频模型。
文本-图像到视频(TI2V)生成是利用语义与视觉条件实现可控视频生成的关键问题。现有方法通常通过微调将视觉条件引入文本到视频(T2V)基础模型,成本高且仅支持少数预定义条件设置。为解决这一限制,我们提出一种统一的TI2V生成框架,支持灵活的视觉条件输入。进一步提出一种创新的无训练方法FlexTI2V,可在任意位置、任意数量图像条件下对T2V基础模型进行条件控制。首先将条件图像反演至隐空间中的噪声表示;随后在T2V模型去噪过程中,采用新颖的随机块替换策略,通过局部图像块将视觉特征融入视频表示。为平衡创造性和保真度,引入动态控制机制,按帧调节视觉条件强度。大量实验表明,该方法在显著优于现有无训练图像条件方法的同时,可泛化至基于UNet和Transformer的模型架构。
原文摘要 · Abstract (English)
Text-image-to-video (TI2V) generation is a critical problem for controllable video generation using both semantic and visual conditions. Most existing methods typically add visual conditions to text-to-video (T2V) foundation models by finetuning, which is costly in resources and only limited to a few pre-defined conditioning settings. To tackle these constraints, we introduce a unified formulation for TI2V generation with flexible visual conditioning. Furthermore, we propose an innovative training-free approach, dubbed FlexTI2V, that can condition T2V foundation models on an arbitrary amount of images at arbitrary positions. Specifically, we firstly invert the condition images to noisy representation in a latent space. Then, in the denoising process of T2V models, our method uses a novel random patch swapping strategy to incorporate visual features into video representations through local image patches. To balance creativity and fidelity, we use a dynamic control mechanism to adjust the strength of visual conditioning to each video frame. Extensive experiments validate that our method surpasses previous training-free image conditioning methods by a notable margin. Our method can also generalize to both UNet-based and transformer-based architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。