arXiv:2505.17358cs.CV2025-05NeurIPS被引 4

用镜头模糊信息让预训练模型直接输出毫米级深度图。

Repurposing Marigold for Zero-Shot Metric Depth Estimation via Defocus Blur Cues

  • 用双光圈拍摄获取模糊差异,注入模型推理过程
  • 无需训练,在真实数据上实现更准的毫米级深度估计
  • 适合希望快速部署零样本深度估计的开发者

近期单目度量深度估计(MMDE)方法在零样本泛化方面取得进展,但仍存在对分布外数据性能显著下降的问题。本文通过在推理时注入景深模糊线索,将已预训练的扩散模型Marigold转化为无需训练的度量深度预测器。具体做法是从同一视角分别使用小光圈和大光圈拍摄两张图像,利用基于景深模糊成像模型的损失函数,优化度量深度缩放参数与Marigold的噪声隐变量。在自建的真实数据集上,本方法在定量与定性指标上均优于现有最先进零样本MMDE方法。

原文摘要 · Abstract (English)

Recent monocular metric depth estimation (MMDE) methods have made notable progress towards zero-shot generalization. However, they still exhibit a significant performance drop on out-of-distribution datasets. We address this limitation by injecting defocus blur cues at inference time into Marigold, a \textit{pre-trained} diffusion model for zero-shot, scale-invariant monocular depth estimation (MDE). Our method effectively turns Marigold into a metric depth predictor in a training-free manner. To incorporate defocus cues, we capture two images with a small and a large aperture from the same viewpoint. To recover metric depth, we then optimize the metric depth scaling parameters and the noise latents of Marigold at inference time using gradients from a loss function based on the defocus-blur image formation model. We compare our method against existing state-of-the-art zero-shot MMDE methods on a self-collected real dataset, showing quantitative and qualitative improvements.

深度估计扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。