构建1400万对高质量视频文本数据集,免爬取、可直接用。
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
- 融合多源开放视频,统一去重与质量过滤。
- 1400万条数据,长文本描述精准匹配动作与时间结构。
- 适合训练开源视频生成模型,解决数据瓶颈问题。
自Sora发布以来,文本生成视频领域备受关注,但开源模型仍面临数据瓶颈:缺乏大规模、高质量且易于获取的视频-文本语料库。现有公开数据集通常依赖人工爬取YouTube,因链接失效和访问限制导致可用数据量低,且存在版权不确定性。本文提出ViMix-14M,一个约1400万对视频-文本数据的精选多源数据集,提供免爬取、即开即用的访问方式,并配有长篇、高质量、与视频高度对齐的描述。该数据集通过整合多样化的开放视频来源,经统一去重、质量筛选及多粒度、基于真实标注的重标注流程,提升描述对动作、场景和时间结构的匹配度。我们在多模态检索、文本生成视频和视频问答任务中评估该数据集,结果均优于同类数据集。我们希望此项工作能消除训练和微调开源视频基础模型的关键障碍,并为构建高质量、通用性强的视频-文本数据集提供参考。
原文摘要 · Abstract (English)
Text-to-video generation has surged in interest since Sora, yet open-source models still face a data bottleneck: there is no large, high-quality, easily obtainable video-text corpus. Existing public datasets typically require manual YouTube crawling, which yields low usable volume due to link rot and access limits, and raises licensing uncertainty. This work addresses this challenge by introducing ViMix-14M, a curated multi-source video-text dataset of around 14 million pairs that provides crawl-free, download-ready access and long-form, high-quality captions tightly aligned to video. ViMix-14M is built by merging diverse open video sources, followed by unified de-duplication and quality filtering, and a multi-granularity, ground-truth-guided re-captioning pipeline that refines descriptions to better match actions, scenes, and temporal structure. We evaluate the dataset by multimodal retrieval, text-to-video generation, and video question answering tasks, observing consistent improvements over counterpart datasets. We hope this work can help removing the key barrier to training and fine-tuning open-source video foundation models, and provide insights of building high-quality and generalizable video-text datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。