arXiv:2501.11841cs.CV2025-01综述被引 52

解决单目深度估计无尺度问题,实现真实世界尺寸的精准深度预测。

Survey on Monocular Metric Depth Estimation

  • 结合几何约束与深度网络,从单张图像恢复带绝对尺度的深度图。
  • 在KITTI、NYU-Depth v2等数据集上实现亚米级精度,显著提升3D重建可靠性。
  • 适合自动驾驶、机器人导航等需精确空间感知的场景应用。

单目深度估计(MDE)支持空间理解、三维重建和自主导航,但现有深度学习方法通常仅输出相对深度,缺乏一致的度量尺度,限制了其在视觉SLAM、精确三维建模和视图合成中的可靠性。单目度量深度估计(MMDE)通过生成具有绝对尺度的深度图,确保几何一致性,无需额外标定即可部署。本综述系统回顾了MMDE的发展历程,涵盖基于几何的方法到最新深度模型,重点分析推动进展的关键数据集:KITTI、NYU-Depth v2、ApolloScape和TartanAir,评估其模态、场景类型与应用领域。方法论进展包括领域泛化、边界保持及合成与真实数据融合。评估了无监督/半监督学习、基于块推理、架构创新与生成建模的技术优劣。通过整合当前成果,强调高质量数据集的重要性,并指明开放挑战,为推进MMDE提供结构化参考,支持其在实际计算机视觉系统中的应用。

原文摘要 · Abstract (English)

Monocular Depth Estimation (MDE) enables spatial understanding, 3D reconstruction, and autonomous navigation, yet deep learning approaches often predict only relative depth without a consistent metric scale. This limitation reduces reliability in applications such as visual SLAM, precise 3D modeling, and view synthesis. Monocular Metric Depth Estimation (MMDE) overcomes this challenge by producing depth maps with absolute scale, ensuring geometric consistency and enabling deployment without additional calibration. This survey reviews the evolution of MMDE, from geometry-based methods to state-of-the-art deep models, with emphasis on the datasets that drive progress. Key benchmarks, including KITTI, NYU-D, ApolloScape, and TartanAir, are examined in terms of modality, scene type, and application domain. Methodological advances are analyzed, covering domain generalization, boundary preservation, and the integration of synthetic and real data. Techniques such as unsupervised and semi-supervised learning, patch-based inference, architectural innovations, and generative modeling are evaluated for their strengths and limitations. By synthesizing current progress, highlighting the importance of high-quality datasets, and identifying open challenges, this survey provides a structured reference for advancing MMDE and supporting its adoption in real-world computer vision systems.

深度估计单目视觉度量尺度三维重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。