arXiv:2510.18291cs.CV2025-10ICCV

用立体视觉引导扩散模型,实现无需重训练的精准深度估计

GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation

  • 将深度估计重构为逆问题,结合图像与立体几何约束
  • 在复杂场景中优于当前最优方法,尤其对透明/反光表面表现优异
  • 无需重新训练,可直接集成到现有单目深度模型中

我们提出一种新型框架,用于度量深度估计,通过引入立体视觉引导来增强预训练的基于扩散的单目深度估计(DB-MDE)模型。现有DB-MDE方法在预测相对深度方面表现优秀,但在单图场景下因尺度模糊性,难以准确估计绝对度量深度。为此,我们将深度估计重构为逆问题,利用预训练的潜在扩散模型(LDMs),在给定RGB图像条件下,结合基于立体的几何约束,学习尺度与偏移以实现精确深度恢复。该方法无需训练,可无缝集成至现有DB-MDE框架,并在室内、室外及复杂环境中具有良好泛化能力。大量实验表明,本方法在挑战性场景(如半透明和镜面表面)中达到或超越当前最先进水平,且无需重新训练。

原文摘要 · Abstract (English)

We introduce a novel framework for metric depth estimation that enhances pretrained diffusion-based monocular depth estimation (DB-MDE) models with stereo vision guidance. While existing DB-MDE methods excel at predicting relative depth, estimating absolute metric depth remains challenging due to scale ambiguities in single-image scenarios. To address this, we reframe depth estimation as an inverse problem, leveraging pretrained latent diffusion models (LDMs) conditioned on RGB images, combined with stereo-based geometric constraints, to learn scale and shift for accurate depth recovery. Our training-free solution seamlessly integrates into existing DB-MDE frameworks and generalizes across indoor, outdoor, and complex environments. Extensive experiments demonstrate that our approach matches or surpasses state-of-the-art methods, particularly in challenging scenarios involving translucent and specular surfaces, all without requiring retraining.

深度估计扩散模型立体视觉无训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。