构建首个200万条长镜头视频数据集,支持动态长视频生成研究
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
- 通过多维度质量评估筛选超过10秒无剪辑长镜头视频
- 开发分层字幕生成管道,实现每段视频的密集时间标注
- 数据集含丰富运动与多样内容,适合训练长视频生成模型
视频生成模型的效果高度依赖于训练数据的质量。以往多数模型基于短片段视频训练,而近期对直接在长视频上训练长视频生成模型的兴趣日益增长。然而,高质量长视频数据的缺乏阻碍了该方向的发展。为推动长视频生成研究,我们提出一个具备四项关键特征的新数据集:(1)时长至少10秒的长视频;(2)无剪辑的长镜头视频;(3)大范围运动与多样化内容;(4)时间密集的字幕标注。为此,我们设计了一套视频质量评估指标,包括场景切换、动态程度和语义质量,用于从海量源视频中筛选高质量长镜头视频。随后,我们构建了分层视频字幕生成流程,对长视频进行时间密集标注。基于此流程,我们创建了首个长镜头视频数据集LVD-2M,包含200万条超过10秒的长镜头视频,并配有密集时间字幕。我们通过微调视频生成模型验证了该数据集在生成具有动态运动的长视频上的有效性。我们相信本工作将显著推动未来长视频生成的研究。
原文摘要 · Abstract (English)
The efficacy of video generation models heavily depends on the quality of their training datasets. Most previous video generation models are trained on short video clips, while recently there has been increasing interest in training long video generation models directly on longer videos. However, the lack of such high-quality long videos impedes the advancement of long video generation. To promote research in long video generation, we desire a new dataset with four key features essential for training long video generation models: (1) long videos covering at least 10 seconds, (2) long-take videos without cuts, (3) large motion and diverse contents, and (4) temporally dense captions. To achieve this, we introduce a new pipeline for selecting high-quality long-take videos and generating temporally dense captions. Specifically, we define a set of metrics to quantitatively assess video quality including scene cuts, dynamic degrees, and semantic-level quality, enabling us to filter high-quality long-take videos from a large amount of source videos. Subsequently, we develop a hierarchical video captioning pipeline to annotate long videos with temporally-dense captions. With this pipeline, we curate the first long-take video dataset, LVD-2M, comprising 2 million long-take videos, each covering more than 10 seconds and annotated with temporally dense captions. We further validate the effectiveness of LVD-2M by fine-tuning video generation models to generate long videos with dynamic motions. We believe our work will significantly contribute to future research in long video generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。