用立体视觉引导扩散模型,实现无需重训练的精准深度估计
GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
- 将深度估计重构为逆问题,结合图像与立体几何约束
- 在复杂场景中优于当前最优方法,尤其对透明/反光表面表现优异
- 无需重新训练,可直接集成到现有单目深度模型中
我们提出一种新型框架,用于度量深度估计,通过引入立体视觉引导来增强预训练的基于扩散的单目深度估计(DB-MDE)模型。现有DB-MDE方法在预测相对深度方面表现优秀,但在单图场景下因尺度模糊性,难以准确估计绝对度量深度。为此,我们将深度估计重构为逆问题,利用预训练的潜在扩散模型(LDMs),在给定RGB图像条件下,结合基于立体的几何约束,学习尺度与偏移以实现精确深度恢复。该方法无需训练,可无缝集成至现有DB-MDE框架,并在室内、室外及复杂环境中具有良好泛化能力。大量实验表明,本方法在挑战性场景(如半透明和镜面表面)中达到或超越当前最先进水平,且无需重新训练。
原文摘要 · Abstract (English)
We introduce a novel framework for metric depth estimation that enhances pretrained diffusion-based monocular depth estimation (DB-MDE) models with stereo vision guidance. While existing DB-MDE methods excel at predicting relative depth, estimating absolute metric depth remains challenging due to scale ambiguities in single-image scenarios. To address this, we reframe depth estimation as an inverse problem, leveraging pretrained latent diffusion models (LDMs) conditioned on RGB images, combined with stereo-based geometric constraints, to learn scale and shift for accurate depth recovery. Our training-free solution seamlessly integrates into existing DB-MDE frameworks and generalizes across indoor, outdoor, and complex environments. Extensive experiments demonstrate that our approach matches or surpasses state-of-the-art methods, particularly in challenging scenarios involving translucent and specular surfaces, all without requiring retraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。