用生成模型先验提升单图度量深度估计精度
MetricGold: Leveraging Text-To-Image Latent Diffusion Models for Metric Depth Estimation
- 利用隐空间扩散模型的场景先验增强深度预测
- 在多个数据集上实现更清晰、更高精度的度量深度估计
- 适合关注单目深度与生成模型融合的研究者
从单张图像恢复度量深度仍是计算机视觉中的基础挑战,需同时具备场景理解与准确尺度还原能力。尽管深度学习已推动单目深度估计发展,现有模型在未知场景和零样本情形下仍表现不佳,尤其在尺度不变的度量深度预测方面。我们提出MetricGold,一种利用生成式扩散模型丰富先验来提升度量深度估计的新方法。基于MariGold、DDVM和Depth Anything V2的最新进展,该方法结合了隐空间扩散模型、对数尺度度量深度表示以及合成数据训练。仅用一张RTX 3090显卡,在两天内完成高效训练,使用HyperSIM、VirtualKitti和TartanAir生成的逼真合成数据。实验表明,该方法在多个数据集上具有稳健泛化能力,相比现有方法生成更锐利、质量更高的度量深度图。
原文摘要 · Abstract (English)
Recovering metric depth from a single image remains a fundamental challenge in computer vision, requiring both scene understanding and accurate scaling. While deep learning has advanced monocular depth estimation, current models often struggle with unfamiliar scenes and layouts, particularly in zero-shot scenarios and when predicting scale-ergodic metric depth. We present MetricGold, a novel approach that harnesses generative diffusion model's rich priors to improve metric depth estimation. Building upon recent advances in MariGold, DDVM and Depth Anything V2 respectively, our method combines latent diffusion, log-scaled metric depth representation, and synthetic data training. MetricGold achieves efficient training on a single RTX 3090 within two days using photo-realistic synthetic data from HyperSIM, VirtualKitti, and TartanAir. Our experiments demonstrate robust generalization across diverse datasets, producing sharper and higher quality metric depth estimates compared to existing approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。