用20万美金训练出顶级视频生成模型,开源可复现。
Open-Sora 2.0: Training a Commercial-Level Video Generation Model in $200k

- 通过数据筛选、架构优化与系统调优,大幅降低训练成本。
- 人类评估和VBench评分达到领先水平,媲美闭源模型。
- 全开源发布,适合研究者与创作者快速部署与创新。
过去一年中,视频生成模型取得显著进展,但通常伴随模型规模增大、数据量上升及训练算力需求激增。本文介绍Open-Sora 2.0——仅花费20万美元即训练完成的商业级视频生成模型。该成果表明,顶尖视频生成模型的训练成本可被有效控制。我们详细阐述了实现这一效率突破的关键技术,包括数据筛选、模型架构、训练策略与系统优化。根据人类评估结果及VBench得分,Open-Sora 2.0性能可与全球领先的视频生成模型相媲美,包括开源的HunyuanVideo和闭源的Runway Gen-3 Alpha。通过将Open-Sora 2.0完全开源,我们致力于推动先进视频生成技术的普及,促进内容创作领域的广泛创新。所有资源已公开于:https://github.com/hpcaitech/Open-Sora。
原文摘要 · Abstract (English)
Video generation models have achieved remarkable progress in the past year. The quality of AI video continues to improve, but at the cost of larger model size, increased data quantity, and greater demand for training compute. In this report, we present Open-Sora 2.0, a commercial-level video generation model trained for only $200k. With this model, we demonstrate that the cost of training a top-performing video generation model is highly controllable. We detail all techniques that contribute to this efficiency breakthrough, including data curation, model architecture, training strategy, and system optimization. According to human evaluation results and VBench scores, Open-Sora 2.0 is comparable to global leading video generation models including the open-source HunyuanVideo and the closed-source Runway Gen-3 Alpha. By making Open-Sora 2.0 fully open-source, we aim to democratize access to advanced video generation technology, fostering broader innovation and creativity in content creation. All resources are publicly available at: https://github.com/hpcaitech/Open-Sora.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。