Goku用流模型实现图文生成新标杆,视频质量领先。
Goku: Flow Based Video Generative Foundation Models
- 基于修正流Transformer构建统一图文生成框架
- 文本到图像生成在GenEval达0.76,视频任务VBench达84.85
- 适合关注多模态生成与大规模训练的开发者
本文提出Goku,一个基于修正流Transformer的先进联合图像与视频生成模型家族,实现了行业领先性能。通过精心设计的数据清洗流程、模型架构、流公式和高效训练基础设施,Goku在定性和定量评估中均表现卓越,创下多个任务新纪录。具体而言,其文本到图像生成在GenEval上得分为0.76,在DPG-Bench上达83.65;文本到视频生成在VBench上达到84.85。本工作为构建统一图文生成模型提供了宝贵实践与技术推进。
原文摘要 · Abstract (English)
This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality visual generation, including the data curation pipeline, model architecture design, flow formulation, and advanced infrastructure for efficient and robust large-scale training. The Goku models demonstrate superior performance in both qualitative and quantitative evaluations, setting new benchmarks across major tasks. Specifically, Goku achieves 0.76 on GenEval and 83.65 on DPG-Bench for text-to-image generation, and 84.85 on VBench for text-to-video tasks. We believe that this work provides valuable insights and practical advancements for the research community in developing joint image-and-video generation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。