arXiv:2409.09896cs.CVcs.LG2024-09被引 6

用稀疏数据训练扩散模型,实现单图像素级精确深度估计

GRIN: Zero-Shot Metric Depth with Pixel-Level Diffusion

  • 用3D位置编码引导扩散过程,结合全局与局部图像特征
  • 在8个场景上实现零样本度量深度新纪录,无需预训练
  • 适合需要低标注成本的3D重建与自动驾驶应用

从单张图像进行3D重建是计算机视觉中的长期难题。基于学习的方法通过利用大规模有标签和无标签数据来缓解尺度模糊性,生成跨域准确的几何先验。当前最先进方法在零样本相对与度量深度估计上表现优异。最近,扩散模型在表示学习中展现出卓越的可扩展性和泛化能力。然而,这些模型沿用原本用于图像生成的工具,仅能处理密集真值,而现实中深度标签往往稀疏且无结构。本文提出GRIN,一种高效扩散模型,可直接处理稀疏非结构化训练数据。通过融合图像特征与3D几何位置编码,实现对扩散过程的全局与局部条件控制,生成像素级深度预测。在八个室内外数据集上的综合实验表明,即使从零开始训练,GRIN仍建立零样本单目度量深度估计的新基准。

原文摘要 · Abstract (English)

3D reconstruction from a single image is a long-standing problem in computer vision. Learning-based methods address its inherent scale ambiguity by leveraging increasingly large labeled and unlabeled datasets, to produce geometric priors capable of generating accurate predictions across domains. As a result, state of the art approaches show impressive performance in zero-shot relative and metric depth estimation. Recently, diffusion models have exhibited remarkable scalability and generalizable properties in their learned representations. However, because these models repurpose tools originally designed for image generation, they can only operate on dense ground-truth, which is not available for most depth labels, especially in real-world settings. In this paper we present GRIN, an efficient diffusion model designed to ingest sparse unstructured training data. We use image features with 3D geometric positional encodings to condition the diffusion process both globally and locally, generating depth predictions at a pixel-level. With comprehensive experiments across eight indoor and outdoor datasets, we show that GRIN establishes a new state of the art in zero-shot metric monocular depth estimation even when trained from scratch.

深度估计扩散模型单目重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。