arXiv:2501.05147cs.CVcs.AI2025-01综述被引 6

系统梳理深度学习在视觉深度估计中的研究进展,覆盖主流方法与数据挑战。

A Systematic Literature Review on Deep Learning-based Depth Estimation in Computer Vision

  • 按单目、双目、多视角分类整理深度估计方法
  • 发现KITTI等20个数据集被广泛使用,35种基础模型中ResNet系列最常见
  • 指出真实深度标注数据稀缺是当前核心难题,适合领域研究者参考

深度估计(DE)提供场景空间信息,支持三维重建、目标检测与场景理解等任务。传统方法依赖手工特征,泛化能力差且需大量调参。深度学习方法可自动提取特征,适应多样场景并具有良好泛化性。为系统总结最新进展,本研究对1284篇文献进行筛选,最终纳入59篇高质量原始研究。分析显示,深度学习方法主要应用于单目、双目和多视角深度估计;共使用20个公开数据集,其中KITTI、NYU Depth V2和Make 3D使用最多;采用29种评估指标,报告了35种基础模型,前五名分别为ResNet-50、ResNet-18、ResNet-101、U-Net和VGG-16。研究指出,缺乏真实深度标注数据是当前最主要的挑战。

原文摘要 · Abstract (English)

Depth estimation (DE) provides spatial information about a scene and enables tasks such as 3D reconstruction, object detection, and scene understanding. Recently, there has been an increasing interest in using deep learning (DL)-based methods for DE. Traditional techniques rely on handcrafted features that often struggle to generalise to diverse scenes and require extensive manual tuning. However, DL models for DE can automatically extract relevant features from input data, adapt to various scene conditions, and generalise well to unseen environments. Numerous DL-based methods have been developed, making it necessary to survey and synthesize the state-of-the-art (SOTA). Previous reviews on DE have mainly focused on either monocular or stereo-based techniques, rather than comprehensively reviewing DE. Furthermore, to the best of our knowledge, there is no systematic literature review (SLR) that comprehensively focuses on DE. Therefore, this SLR study is being conducted. Initially, electronic databases were searched for relevant publications, resulting in 1284 publications. Using defined exclusion and quality criteria, 128 publications were shortlisted and further filtered to select 59 high-quality primary studies. These studies were analysed to extract data and answer defined research questions. Based on the results, DL methods were developed for mainly three different types of DE: monocular, stereo, and multi-view. 20 publicly available datasets were used to train, test, and evaluate DL models for DE, with KITTI, NYU Depth V2, and Make 3D being the most used datasets. 29 evaluation metrics were used to assess the performance of DE. 35 base models were reported in the primary studies, and the top five most-used base models were ResNet-50, ResNet-18, ResNet-101, U-Net, and VGG-16. Finally, the lack of ground truth data was among the most significant challenges reported by primary studies.

深度估计深度学习文献综述计算机视觉

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。