构建首个融合动态镜头与剧情连贯性的百万级视频数据集,推动长时一致性视频生成
DropletVideo: A Dataset and Approach to Explore Integral Spatio-Temporal Consistent Video Generation
- 提出整合镜头运动与剧情发展的整体时空一致性机制
- 构建含1000万条带206词平均描述的动态视频数据集
- 适合视频生成、影视制作及多视角场景建模研究者使用
时空一致性是视频生成的关键挑战。高质量视频需在不同视角下保持物体与场景视觉一致,并确保剧情合理连贯。现有研究多聚焦于时间或空间一致性,或其简单组合,如仅在提示后添加镜头移动描述,但未约束移动带来的画面变化。镜头运动可能引入新物体或移除原有元素,从而干扰先前叙事,尤其在频繁镜头切换的视频中,多重剧情交互愈发复杂。本文首次系统探讨整体时空一致性,综合考虑剧情演进与镜头技术的协同作用,以及前期内容对后续生成的长期影响。研究涵盖数据集构建至模型开发:首先构建包含1000万条视频的DropletVideo-10M数据集,每段视频均配有平均206词的详细标注,涵盖镜头运动与剧情发展;随后训练并验证DropletVideo模型,在视频生成中显著提升时空一致性表现。数据集与模型已公开于https://dropletx.github.io。
原文摘要 · Abstract (English)
Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying viewpoints. Prior research, especially in open-source projects, primarily focuses on either temporal or spatial consistency, or their basic combination, such as appending a description of a camera movement after a prompt without constraining the outcomes of this movement. However, camera movement may introduce new objects to the scene or eliminate existing ones, thereby overlaying and affecting the preceding narrative. Especially in videos with numerous camera movements, the interplay between multiple plots becomes increasingly complex. This paper introduces and examines integral spatio-temporal consistency, considering the synergy between plot progression and camera techniques, and the long-term impact of prior content on subsequent generation. Our research encompasses dataset construction through to the development of the model. Initially, we constructed a DropletVideo-10M dataset, which comprises 10 million videos featuring dynamic camera motion and object actions. Each video is annotated with an average caption of 206 words, detailing various camera movements and plot developments. Following this, we developed and trained the DropletVideo model, which excels in preserving spatio-temporal coherence during video generation. The DropletVideo dataset and model are accessible at https://dropletx.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。