一模型同时生成视频、分割图和深度图,提升画面连贯性。
Unified Dense Prediction of Video Diffusion
- 用颜色映射融合分割与深度信息,统一生成视频
- 在不增加计算量下显著提升运动流畅度和一致性
- 首个包含文本、视频、分割和深度图的大规模数据集
我们提出一种统一网络,可从文本提示中同时生成视频及其对应的实体分割图和深度图。采用颜色映射表示实体掩码和深度图,将密集预测任务与RGB视频生成紧密结合。引入密集预测信息可在不增加计算成本的前提下提升视频生成的一致性和运动平滑性。通过可学习的任务嵌入,将多个密集预测任务集成到单一模型中,增强灵活性并进一步提升性能。我们还构建了一个大规模密集预测视频数据集~\datasetname,解决了现有数据集无法同时提供文本描述、视频、分割图和深度图的问题。全面实验表明,该方法在视频质量、一致性和运动平滑性方面均优于当前最先进水平,且具有高效率。
原文摘要 · Abstract (English)
We present a unified network for simultaneously generating videos and their corresponding entity segmentation and depth maps from text prompts. We utilize colormap to represent entity masks and depth maps, tightly integrating dense prediction with RGB video generation. Introducing dense prediction information improves video generation's consistency and motion smoothness without increasing computational costs. Incorporating learnable task embeddings brings multiple dense prediction tasks into a single model, enhancing flexibility and further boosting performance. We further propose a large-scale dense prediction video dataset~\datasetname, addressing the issue that existing datasets do not concurrently contain captions, videos, segmentation, or depth maps. Comprehensive experiments demonstrate the high efficiency of our method, surpassing the state-of-the-art in terms of video quality, consistency, and motion smoothness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。