开源1000万小时视频数据集,支持多模态预训练
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training
- 从CommonCrawl抓取13亿视频链接,下载8000万段视频
- 训练模型在视频-文本、音频-文本任务上表现优异
- 可用视频帧生成图像-文本数据,提升图文检索性能
我们提出LAION-BVD,一个大规模开放视频数据集,用于多模态学习。该数据集包含从CommonCrawl收集的13亿个平台特定视频链接,从中下载了8000万段视频,总时长达1000万小时。数据集设计用于跨视频、音频和图像模态的多模态预训练。通过内容感知场景检测,提取片段并合成视频与音频描述。基于这些数据训练的模型在标准视频-文本和音频-文本基准测试中表现优异,且随着训练规模或模型规模增大持续提升。此外,我们探索以视频帧作为图像-文本数据源,提取场景切换帧。这些帧的视觉分布显著区别于标准网页图像语料库,使用该数据训练的模型在图像-文本检索任务中表现强劲。我们已向研究社区发布LAION-BVD,极大拓展了开放获取多模态视频数据的规模。
原文摘要 · Abstract (English)
We present LAION-BVD, a large-scale open video dataset for multimodal learning, which contains 1.3B platform-specific video URLs collected from CommonCrawl. From these, we download 80M videos with a total duration of 10 million hours. The dataset is designed for multimodal pre-training across the video, audio, and image modalities. Using content-aware scene detection, we extract clips for which we synthetically generate video and audio captions. Models trained on these data achieve competitive performance on standard video-text and audio-text benchmarks, with consistent improvements as training or model scale increases. Additionally, we explore video frames as an alternative source of image-text data by extracting scene-changing frames. These frames exhibit a visual distribution distinct from standard web image corpora, and models trained on this dataset achieve strong image-text retrieval performance. We release LAION-BVD to the research community. It significantly expands open access to multimodal videos at an unprecedented scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。