8B参数视频生成模型,4周训练即达顶尖效果。
ContentV: Efficient Training of Video Generation Models with Limited Compute
- 用预训练图像模型复用架构,降低视频生成复杂度。
- 四週訓練在256塊64GB NPU上達85.14分(VBench)。
- 無需額外人工標註,提升生成質量的強化學習框架。
近期視頻生成技術發展迅速,但計算成本持續攀升。本文提出內容為80億參數的文本到視頻模型ContentV,僅用256塊64GB神經處理單元(NPUs)訓練四週,即在VBench評分達85.14,達到當前領先水平。ContentV能根據文本提示生成多分辨率、多時長的高質量、多樣化視頻,其核心創新包括:(1)極簡架構,最大化重用預訓練圖像生成模型;(2)系統化的多階段訓練策略,採用流匹配提升效率;(3)低成本人類反饋強化學習框架,提升生成品質且無需額外人工標註。所有代碼與模型均公開於https://contentv.github.io。
原文摘要 · Abstract (English)
Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter text-to-video model that achieves state-of-the-art performance (85.14 on VBench) after training on 256 x 64GB Neural Processing Units (NPUs) for merely four weeks. ContentV generates diverse, high-quality videos across multiple resolutions and durations from text prompts, enabled by three key innovations: (1) A minimalist architecture that maximizes reuse of pre-trained image generation models for video generation; (2) A systematic multi-stage training strategy leveraging flow matching for enhanced efficiency; and (3) A cost-effective reinforcement learning with human feedback framework that improves generation quality without requiring additional human annotations. All the code and models are available at: https://contentv.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。