300亿参数模型,图文生成102帧视频,性能业界领先
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
- 基于图文输入的分步生成架构,支持长视频连续输出
- 在自建基准上超越开源与商用模型,102帧生成质量最优
- 适合需要高保真视频生成的研究与工业应用
我们提出 Step-Video-TI2V,一个拥有 300 亿参数的前沿文本驱动图像到视频生成模型,可基于文本和图像输入生成长达 102 帧的视频。我们构建了新的评估基准 Step-Video-TI2V-Eval,用于评测文本驱动图像到视频任务,并在此数据集上对比了 Step-Video-TI2V 与开源及商用 TI2V 引擎的表现。实验结果表明,Step-Video-TI2V 在图像到视频生成任务中达到当前最佳性能。相关代码与模型已公开于 https://github.com/stepfun-ai/Step-Video-TI2V。
原文摘要 · Abstract (English)
We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and image inputs. We build Step-Video-TI2V-Eval as a new benchmark for the text-driven image-to-video task and compare Step-Video-TI2V with open-source and commercial TI2V engines using this dataset. Experimental results demonstrate the state-of-the-art performance of Step-Video-TI2V in the image-to-video generation task. Both Step-Video-TI2V and Step-Video-TI2V-Eval are available at https://github.com/stepfun-ai/Step-Video-TI2V.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。