arXiv:2409.02095cs.CVcs.AI2024-09CVPR被引 232

无需相机参数即可生成长达110帧的精准视频深度图

DepthCrafter: Generating Consistent Long Depth Sequences for Open-world Videos

  • 基于扩散模型三阶段训练,实现无额外信息的视频深度生成
  • 单次输出最长110帧深度序列,保持时序一致性与细节精度
  • 适用于长视频处理,适合影视特效与条件视频生成场景

在开放世界视频中估计深度极具挑战性,因其外观、内容、运动和镜头变化多样。我们提出DepthCrafter,一种无需相机位姿或光流等额外信息的创新方法,可生成具有精细细节和时序一致性的长视频深度序列。通过从预训练图像到视频扩散模型出发,采用精心设计的三阶段训练策略,使模型具备对开放世界视频的强泛化能力。该方法支持一次性生成任意长度的深度序列,最长可达110帧,并从真实与合成数据集中提取高精度深度信息与丰富内容多样性。此外,我们提出分段估计与无缝拼接的推理策略,可处理超长视频。在多个数据集上的全面评估表明,DepthCrafter在零样本设置下达到当前最优性能。同时,该方法可赋能深度引导的视觉特效与条件视频生成等下游应用。

原文摘要 · Abstract (English)

Estimating video depth in open-world scenarios is challenging due to the diversity of videos in appearance, content motion, camera movement, and length. We present DepthCrafter, an innovative method for generating temporally consistent long depth sequences with intricate details for open-world videos, without requiring any supplementary information such as camera poses or optical flow. The generalization ability to open-world videos is achieved by training the video-to-depth model from a pre-trained image-to-video diffusion model, through our meticulously designed three-stage training strategy. Our training approach enables the model to generate depth sequences with variable lengths at one time, up to 110 frames, and harvest both precise depth details and rich content diversity from realistic and synthetic datasets. We also propose an inference strategy that can process extremely long videos through segment-wise estimation and seamless stitching. Comprehensive evaluations on multiple datasets reveal that DepthCrafter achieves state-of-the-art performance in open-world video depth estimation under zero-shot settings. Furthermore, DepthCrafter facilitates various downstream applications, including depth-based visual effects and conditional video generation.

视频深度扩散模型长序列生成无监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。