arXiv:2506.20756cs.CV2025-06中稿 · any journal or con…被引 5

用立体匹配+扩散模型,提升视频深度估计的准确与一致

StereoDiff: Stereo-Diffusion Synergy for Video Depth Estimation

  • 分两阶段:静态区用立体匹配,动态区用视频扩散模型
  • 在真实场景下零样本测试中达到当前最佳性能
  • 适合需要高一致性视频深度的应用,如AR/VR

现有视频深度估计方法多沿用图像深度估计范式,通过大规模数据微调预训练视频扩散模型。但本文指出,视频深度估计并非图像深度的简单延伸——动态与静态区域的时间一致性要求本质不同。静态区域(如背景)可通过跨帧立体匹配获得更强全局3D线索,实现更优一致性;而动态区域因违反三角化约束,仍需依赖大规模视频深度数据学习平滑过渡。基于此,我们提出StereoDiff,一种两阶段视频深度估计算法:在静态区域融合立体匹配,在动态区域采用视频深度扩散模型。通过频域分析证明二者在数学上互补,协同提升性能。在零样本、真实世界、动态视频深度基准测试中(涵盖室内外场景),StereoDiff展现出最优表现,显著提升深度估计的一致性与准确性。

原文摘要 · Abstract (English)

Recent video depth estimation methods achieve great performance by following the paradigm of image depth estimation, i.e., typically fine-tuning pre-trained video diffusion models with massive data. However, we argue that video depth estimation is not a naive extension of image depth estimation. The temporal consistency requirements for dynamic and static regions in videos are fundamentally different. Consistent video depth in static regions, typically backgrounds, can be more effectively achieved via stereo matching across all frames, which provides much stronger global 3D cues. While the consistency for dynamic regions still should be learned from large-scale video depth data to ensure smooth transitions, due to the violation of triangulation constraints. Based on these insights, we introduce StereoDiff, a two-stage video depth estimator that synergizes stereo matching for mainly the static areas with video depth diffusion for maintaining consistent depth transitions in dynamic areas. We mathematically demonstrate how stereo matching and video depth diffusion offer complementary strengths through frequency domain analysis, highlighting the effectiveness of their synergy in capturing the advantages of both. Experimental results on zero-shot, real-world, dynamic video depth benchmarks, both indoor and outdoor, demonstrate StereoDiff's SoTA performance, showcasing its superior consistency and accuracy in video depth estimation.

视频深度估计立体匹配扩散模型一致性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。