arXiv:2410.02073cs.CVcs.LG2024-10被引 553

单目深度估计模型可在0.3秒内生成高精度、带真实尺度的清晰深度图。

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

  • 基于多尺度视觉变换器,实现密集预测与高效推理。
  • 225万像素深度图仅需0.3秒,绝对尺度准确且无需相机参数。
  • 支持零样本迁移,适合实时应用与无标定场景。

我们提出一种用于零样本度量单目深度估计的基础模型——Depth Pro。该模型可生成高分辨率、边缘锐利且富含高频细节的深度图。其预测结果具备真实尺度,不依赖相机内参等元数据。模型推理速度极快,在标准GPU上仅需0.3秒即可生成225万像素的深度图。技术贡献包括:高效的多尺度视觉变换器用于密集预测;结合真实与合成数据的训练策略,实现高精度度量估计与精细边界追踪;专为深度图边界精度设计的评估指标;以及当前最优的单图焦距估计能力。大量实验验证了各项设计的有效性,表明Depth Pro在多个维度上优于现有方法。代码与权重已开源:https://github.com/apple/ml-depth-pro

原文摘要 · Abstract (English)

We present a foundation model for zero-shot metric monocular depth estimation. Our model, Depth Pro, synthesizes high-resolution depth maps with unparalleled sharpness and high-frequency details. The predictions are metric, with absolute scale, without relying on the availability of metadata such as camera intrinsics. And the model is fast, producing a 2.25-megapixel depth map in 0.3 seconds on a standard GPU. These characteristics are enabled by a number of technical contributions, including an efficient multi-scale vision transformer for dense prediction, a training protocol that combines real and synthetic datasets to achieve high metric accuracy alongside fine boundary tracing, dedicated evaluation metrics for boundary accuracy in estimated depth maps, and state-of-the-art focal length estimation from a single image. Extensive experiments analyze specific design choices and demonstrate that Depth Pro outperforms prior work along multiple dimensions. We release code and weights at https://github.com/apple/ml-depth-pro

单目深度视觉变换器实时推理度量估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。