构建3600万条高质量视频数据,提升文本与视频内容的精细一致性。
Koala-36M: A Large-scale Video Dataset Improving Consistency between Fine-grained Conditions and Video Content
- 用概率分布线性分类器精准检测视频切换点,增强时间连贯性。
- 生成平均200词的结构化字幕,显著提升图文对齐效果。
- 设计多指标融合的视频质量评分系统,筛选优质视频素材。
随着视觉生成技术的快速发展,视频数据集规模呈指数增长。数据集质量对视频生成模型性能至关重要。我们指出时间切分精度、详细字幕和视频质量过滤是决定数据集质量的三大关键因素。然而现有数据集在这些方面存在明显不足。为此,我们提出Koala-36M,一个大规模、高质视频数据集,具有精确的时间切分、详细的结构化字幕和优良的视频质量。核心在于提升细粒度条件与视频内容的一致性。具体而言,采用基于概率分布的线性分类器提升切换点检测准确率,保障时间一致性;为分割后的视频提供平均长度达200词的结构化字幕,强化文本-视频对齐;设计集成多子指标的视频训练适用性评分(VTSS),从原始语料中筛选高质量视频;最后,在生成模型训练过程中引入多项评估指标,进一步优化细粒度条件建模。实验验证了我们数据处理流程的有效性及所提数据集的高质量。相关代码与数据已公开于https://koala36m.github.io/。
原文摘要 · Abstract (English)
With the continuous progress of visual generation technologies, the scale of video datasets has grown exponentially. The quality of these datasets plays a pivotal role in the performance of video generation models. We assert that temporal splitting, detailed captions, and video quality filtering are three crucial determinants of dataset quality. However, existing datasets exhibit various limitations in these areas. To address these challenges, we introduce Koala-36M, a large-scale, high-quality video dataset featuring accurate temporal splitting, detailed captions, and superior video quality. The essence of our approach lies in improving the consistency between fine-grained conditions and video content. Specifically, we employ a linear classifier on probability distributions to enhance the accuracy of transition detection, ensuring better temporal consistency. We then provide structured captions for the splitted videos, with an average length of 200 words, to improve text-video alignment. Additionally, we develop a Video Training Suitability Score (VTSS) that integrates multiple sub-metrics, allowing us to filter high-quality videos from the original corpus. Finally, we incorporate several metrics into the training process of the generation model, further refining the fine-grained conditions. Our experiments demonstrate the effectiveness of our data processing pipeline and the quality of the proposed Koala-36M dataset. Our dataset and code have been released at https://koala36m.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。