构建5000万视频数据集,训练大规模视频基础模型
Summer-22B: A Systematic Approach to Dataset Engineering and Training at Scale for Video Foundation Model
- 用元数据驱动+多阶段过滤构建高质量视频数据集
- 在5000万片段上训练出可运行的视频基础模型
- 适合大规模视频模型研发团队参考
本文介绍从零开始训练视频基础模型Summer-22B的经验。项目从原始视频采集出发,最终训练出一个包含约5000万视频片段的模型。我们提出结合元数据驱动的数据集筛选、多阶段过滤、μP参数化和超球面约束优化的方法,并开发了Lavender Data系统用于数据管理,采用以推理为导向的架构设计。实验发现:数据工程占主要工作量,不同架构差异较小,且μP超参数迁移在几何约束下仍有效。本报告为大规模视频模型训练提供了可复用的工程经验。
原文摘要 · Abstract (English)
We describe our experience training Summer-22B, a video foundation model developed from scratch. This report documents the engineering challenges, design decisions, and lessons learned while scaling from raw footage collection to a functional model trained on approximately 50 million clips. We outline our approach combining metadata-driven dataset curation, multi-stage filtering, $μ$P parameterization, and hypersphere-constrained optimization. We developed the Lavender Data system for dataset management and adopted inference-aware architectural choices. We share observations on what worked in our setting: dataset engineering consumed the majority of effort, architectural variants showed smaller differences than we expected, and $μ$P hyperparameter transfer appeared effective even under geometric constraints. We hope this account proves useful to others undertaking similar projects.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。