arXiv:2511.08036cs.CV2025-11

无需修改模型,利用视觉大模型先验提升单目深度估计效果

WEDepth: Efficient Adaptation of World Knowledge for Monocular Depth Estimation

  • 用视觉大模型作为多层级特征增强器注入先验知识
  • 在NYU-Depth v2和KITTI上达到新SOTA性能
  • 零样本迁移能力强,适合跨场景应用

单目深度估计(MDE)虽应用广泛,但因从单张二维图像重建三维场景的本质病态性而极具挑战。现代视觉基础模型(VFMs)在大规模多样化数据集上预训练,具备出色的全局理解能力,有助于多种视觉任务。近期研究通过微调这些模型显著提升了MDE性能。受此启发,我们提出WEDepth,一种新方法,在不修改模型结构和预训练权重的前提下,有效激发并利用其内在先验。该方法将视觉基础模型作为多层级特征增强器,系统性地在不同表示层次注入先验知识。在NYU-Depth v2和KITTI数据集上的实验表明,WEDepth达到新的最先进水平,性能优于基于扩散模型的方法(后者需多次前向传播)以及在相对深度上预训练的方法。此外,我们还验证了该方法在多种场景下的强零样本迁移能力。

原文摘要 · Abstract (English)

Monocular depth estimation (MDE) has widely applicable but remains highly challenging due to the inherently ill-posed nature of reconstructing 3D scenes from single 2D images. Modern Vision Foundation Models (VFMs), pre-trained on large-scale diverse datasets, exhibit remarkable world understanding capabilities that benefit for various vision tasks. Recent studies have demonstrated significant improvements in MDE through fine-tuning these VFMs. Inspired by these developments, we propose WEDepth, a novel approach that adapts VFMs for MDE without modi-fying their structures and pretrained weights, while effec-tively eliciting and leveraging their inherent priors. Our method employs the VFM as a multi-level feature en-hancer, systematically injecting prior knowledge at differ-ent representation levels. Experiments on NYU-Depth v2 and KITTI datasets show that WEDepth establishes new state-of-the-art (SOTA) performance, achieving competi-tive results compared to both diffusion-based approaches (which require multiple forward passes) and methods pre-trained on relative depth. Furthermore, we demonstrate our method exhibits strong zero-shot transfer capability across diverse scenarios.

深度估计视觉模型零样本迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。