简单设计实现高效图文视频生成,性能超越主流模型。
STIV: Scalable Text and Image Conditioned Video Generation
- 用帧替换融合图像条件,联合文本图像实现无分类器引导。
- 8.7B模型在512分辨率下,文本到视频得分83.1,图像到视频90.1。
- 结构透明可扩展,适用于预测、插帧、多视角等场景。
视频生成领域虽进展显著,但缺乏清晰可复现的构建方法。本文系统研究模型架构、训练策略与数据处理的协同作用,提出一种简单且可扩展的图文条件视频生成方法STIV。该框架通过帧替换将图像条件融入Diffusion Transformer(DiT),并利用联合图像-文本条件的无分类器引导实现文本与图文双重条件生成。此设计使STIV可同时完成文本到视频(T2V)与图文到视频(TI2V)任务。此外,其结构易于扩展至视频预测、帧插值、多视角生成及长视频生成等应用。在T2I、T2V和TI2V任务上进行全面消融实验,结果显示:一个8.7B参数、512分辨率的STIV模型在VBench T2V测试中达到83.1分,超越CogVideoX-5B、Pika、Kling、Gen-3等开源与闭源领先模型;同规模模型在VBench I2V任务上取得90.1的当前最优成绩。本工作提供了一套透明、可扩展的前沿视频生成构建范式,旨在推动未来研究,加速通用可靠视频生成技术的发展。
原文摘要 · Abstract (English)
The field of video generation has made remarkable advancements, yet there remains a pressing need for a clear, systematic recipe that can guide the development of robust and scalable models. In this work, we present a comprehensive study that systematically explores the interplay of model architectures, training recipes, and data curation strategies, culminating in a simple and scalable text-image-conditioned video generation method, named STIV. Our framework integrates image condition into a Diffusion Transformer (DiT) through frame replacement, while incorporating text conditioning via a joint image-text conditional classifier-free guidance. This design enables STIV to perform both text-to-video (T2V) and text-image-to-video (TI2V) tasks simultaneously. Additionally, STIV can be easily extended to various applications, such as video prediction, frame interpolation, multi-view generation, and long video generation, etc. With comprehensive ablation studies on T2I, T2V, and TI2V, STIV demonstrate strong performance, despite its simple design. An 8.7B model with 512 resolution achieves 83.1 on VBench T2V, surpassing both leading open and closed-source models like CogVideoX-5B, Pika, Kling, and Gen-3. The same-sized model also achieves a state-of-the-art result of 90.1 on VBench I2V task at 512 resolution. By providing a transparent and extensible recipe for building cutting-edge video generation models, we aim to empower future research and accelerate progress toward more versatile and reliable video generation solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。