arXiv:2504.08685cs.CVcs.AI2025-04被引 87

70亿参数视频生成模型仅用66.5万小时显卡训练,性能媲美更大模型。

Seaweed-7B: Cost-Effective Training of Video Generation Foundation Model

  • 采用中等规模扩散模型架构,优化训练策略降低资源消耗。
  • 在66.5万H100 GPU小时下达到领先大模型的生成效果。
  • 适合资源有限的研究者快速复现与轻量微调应用。

本技术报告提出一种低成本训练视频生成基础模型的方法。我们构建了一个约70亿参数(7B)的中等规模研究模型——Seaweed-7B,从头开始训练,共使用66.5万小时的H100 GPU算力。尽管计算资源有限,Seaweed-7B仍表现出与当前更大规模视频生成模型相当甚至更优的性能。在资源受限环境下,设计选择尤为关键。本文重点分析了提升中等规模扩散模型性能的关键决策。实证观察显示:(1) Seaweed-7B性能可媲美甚至超越在更大算力上训练的大型模型;(2) 该模型具备强泛化能力,可通过轻量微调或持续训练有效适配多种下游任务。详见项目页:https://seaweed.video/

原文摘要 · Abstract (English)

This technical report presents a cost-efficient strategy for training a video generation foundation model. We present a mid-sized research model with approximately 7 billion parameters (7B) called Seaweed-7B trained from scratch using 665,000 H100 GPU hours. Despite being trained with moderate computational resources, Seaweed-7B demonstrates highly competitive performance compared to contemporary video generation models of much larger size. Design choices are especially crucial in a resource-constrained setting. This technical report highlights the key design decisions that enhance the performance of the medium-sized diffusion model. Empirically, we make two observations: (1) Seaweed-7B achieves performance comparable to, or even surpasses, larger models trained on substantially greater GPU resources, and (2) our model, which exhibits strong generalization ability, can be effectively adapted across a wide range of downstream applications either by lightweight fine-tuning or continue training. See the project page at https://seaweed.video/

视频生成扩散模型低成本训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。