开源视频预训练数据流水线,可一键对比不同数据配方效果
VidaForge: Open Research Infrastructure for Video Pretraining Data Recipes

- 将视频数据流程抽象为五阶段可执行工作流,支持灵活替换决策节点
- 314万片段、6475小时数据集验证:覆盖更广的数据配方提升下游性能
- 适合想研究数据配方如何影响模型表现的研究者快速实验
视频基础模型依赖大规模预训练数据,但其端到端数据管道大多封闭且难以复用。研究者若想探究视频数据配方对模型预训练的影响,往往需自行搭建大量基础设施。我们提出VidaForge,一种开放研究基础设施,将视频数据配方表示为从原始视频到训练数据的五阶段可执行工作流。工作流中的任一决策均可调整以生成替代数据集,同时保留每个样本的生成过程。为演示该研究工作流,我们对比了不同覆盖范围与质量的数据配方在Wan 2.1和V-JEPA 2.1从零开始预训练中的表现。两种学习目标下,覆盖更广的配方均取得最高下游基准得分,而基于损失的评估则偏好不同配方。本研究展示了VidaForge如何将数据配方选择与下游模型性能关联。我们进一步发布VIDAFORGE-3M,包含314万段场景级视频片段,总计6475小时,附有细粒度标注与数据配方研究所需的清理信号。
原文摘要 · Abstract (English)
Video foundation models increasingly rely on large-scale pretraining data, yet the end-to-end data pipelines behind them remain largely closed and difficult to inspect or reuse. Researchers seeking to understand how video data recipes affect model pretraining often need to build substantial infrastructure before testing even a focused hypothesis. We present VIDAFORGE, an open research infrastructure that represents a video data recipe as an executable five-stage workflow from raw videos to training datasets. A decision in this workflow can be varied to construct alternative datasets while preserving how every sample was produced. To demon strate this research workflow, we compare data recipes with different coverage and quality in early from-scratch pretraining of Wan 2.1 and V-JEPA 2.1. Across both learning objectives, the broader-coverage recipe achieves the highest downstream benchmark scores, while loss-based evaluation favors different recipes. This study demonstrates how VidaForge connects data-recipe choices to downstream model performance. We further release VIDAFORGE-3M, containing 3.14 million scene level clips totaling 6,475 hours, with fine-grained annotations and curation signals for video data-recipe research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。