解决长视频深度估计的尺度与几何不一致问题
DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation
- 用扩散模型引导,同步不同视频窗口的深度尺度
- 引入几何约束,确保视频内深度结构一致
- 无需训练,适合长视频深度重建任务
基于扩散模型的视频深度估计方法虽具强大泛化能力,但长视频预测仍具挑战。现有方法将视频分段处理,导致窗口间累积尺度偏差,且仅依赖2D扩散先验,忽略视频深度的内在3D几何结构,造成几何不一致。本文提出DepthSync,一种无需训练的框架,通过扩散引导实现长视频的尺度与几何一致性深度估计。具体地,引入尺度引导以同步各窗口间的深度尺度,几何引导则基于视频深度的固有3D约束,在窗口内强制几何对齐。两者协同作用,引导去噪过程生成一致的深度预测。在多个数据集上的实验验证了该方法在提升深度估计的尺度与几何一致性方面的有效性,尤其适用于长视频。
原文摘要 · Abstract (English)
Diffusion-based video depth estimation methods have achieved remarkable success with strong generalization ability. However, predicting depth for long videos remains challenging. Existing methods typically split videos into overlapping sliding windows, leading to accumulated scale discrepancies across different windows, particularly as the number of windows increases. Additionally, these methods rely solely on 2D diffusion priors, overlooking the inherent 3D geometric structure of video depths, which results in geometrically inconsistent predictions. In this paper, we propose DepthSync, a novel, training-free framework using diffusion guidance to achieve scale- and geometry-consistent depth predictions for long videos. Specifically, we introduce scale guidance to synchronize the depth scale across windows and geometry guidance to enforce geometric alignment within windows based on the inherent 3D constraints in video depths. These two terms work synergistically, steering the denoising process toward consistent depth predictions. Experiments on various datasets validate the effectiveness of our method in producing depth estimates with improved scale and geometry consistency, particularly for long videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。