人工精选20万条高质量UGC视频,提升视频生成模型训练效果
Tiger200K: Manually Curated High Visual Quality Video Dataset from UGC Platform
- 从UGC平台人工筛选高画质视频,确保视觉与美学质量
- 构建20万条时序一致的视频-文本对,支持先进模型微调
- 适合视频生成、数据集构建与人类评估方向的研究者
近期开源文本到视频生成模型的兴起极大推动了研究进展,但其仍依赖专有训练数据,成为关键瓶颈。现有开源数据集如Koala-36M采用算法过滤从早期平台抓取的视频,仍缺乏足够质量以微调先进视频生成模型。本文提出Tiger200K,一个从用户生成内容(UGC)平台手动筛选的高质量视频数据集。通过强调视觉保真度与审美质量,该数据集凸显了人工在数据筛选中的核心作用。我们提供一套简单高效的处理流程,包括镜头边界检测、OCR、边框识别、运动过滤及精细双语标注,构建出20万条高质量、时序一致的视频-文本对,用于微调和优化视频生成架构。该数据集将持续扩展,并作为开源项目发布,以推进视频生成模型的研究与应用。项目页面:https://tinytigerpan.github.io/tiger200k/
原文摘要 · Abstract (English)
The recent surge in open-source text-to-video generation models has significantly energized the research community, yet their dependence on proprietary training datasets remains a key constraint. While existing open datasets like Koala-36M employ algorithmic filtering of web-scraped videos from early platforms, they still lack the quality required for fine-tuning advanced video generation models. We present Tiger200K, a manually curated high visual quality video dataset sourced from User-Generated Content (UGC) platforms. By prioritizing visual fidelity and aesthetic quality, Tiger200K underscores the critical role of human expertise in data curation, and providing high-quality, temporally consistent video-text pairs for fine-tuning and optimizing video generation architectures through a simple but effective pipeline including shot boundary detection, OCR, border detecting, motion filter and fine bilingual caption. The dataset will undergo ongoing expansion and be released as an open-source initiative to advance research and applications in video generative models. Project page: https://tinytigerpan.github.io/tiger200k/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。