arXiv:2503.15905cs.CVcs.AI2025-03NeurIPS被引 26

用扩散模型先验提升单目深度估计的清晰度与泛化能力

Jasmine: Harnessing Diffusion Prior for Self-supervised Depth Estimation

  • 设计混合图像重建代理任务,无需额外标注即可保留扩散模型细节先验
  • 在KITTI上达到当前最佳性能,跨数据集零样本泛化能力强
  • 适合做自监督深度估计、想利用扩散模型视觉先验的研究者

本文提出Jasmine,首个基于稳定扩散(SD)的自监督单目深度估计框架,有效利用SD的视觉先验来提升无监督预测的清晰度与泛化性。以往基于扩散模型的方法均为有监督,因密集预测需高精度标注。而自监督重投影面临遮挡、无纹理区域和光照变化等固有挑战,导致预测模糊且含伪影,严重破坏了SD的潜在先验。为此,我们构建了一种新型混合图像重建代理任务,在无需额外监督的前提下,通过重构图像本身保持了SD模型的细节先验,同时防止深度估计退化。此外,为解决SD的尺度-平移不变性与自监督尺度不变深度估计之间的本质不匹配问题,我们设计了尺度-位移GRU模块,不仅弥合分布差距,还隔离了扩散输出的细粒度纹理免受重投影损失干扰。大量实验表明,Jasmine在KITTI基准上达到当前最佳性能,并在多个数据集上展现出优异的零样本泛化能力。

原文摘要 · Abstract (English)

In this paper, we propose Jasmine, the first Stable Diffusion (SD)-based self-supervised framework for monocular depth estimation, which effectively harnesses SD's visual priors to enhance the sharpness and generalization of unsupervised prediction. Previous SD-based methods are all supervised since adapting diffusion models for dense prediction requires high-precision supervision. In contrast, self-supervised reprojection suffers from inherent challenges (e.g., occlusions, texture-less regions, illumination variance), and the predictions exhibit blurs and artifacts that severely compromise SD's latent priors. To resolve this, we construct a novel surrogate task of hybrid image reconstruction. Without any additional supervision, it preserves the detail priors of SD models by reconstructing the images themselves while preventing depth estimation from degradation. Furthermore, to address the inherent misalignment between SD's scale and shift invariant estimation and self-supervised scale-invariant depth estimation, we build the Scale-Shift GRU. It not only bridges this distribution gap but also isolates the fine-grained texture of SD output against the interference of reprojection loss. Extensive experiments demonstrate that Jasmine achieves SoTA performance on the KITTI benchmark and exhibits superior zero-shot generalization across multiple datasets.

深度估计扩散模型自监督先验利用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。