探索视觉深度估计的通用大模型,突破小数据与泛化难题。
Towards Depth Foundation Model: Recent Trends in Vision-Based Depth Estimation
- 基于大规模数据训练通用深度模型,提升零样本泛化能力。
- 覆盖单目、双目、多视角等多场景,统一建模框架。
- 适合研究者构建鲁棒深度基础模型,推动自动驾驶等领域应用。
深度估计是3D计算机视觉的基础任务,对三维重建、自由视点渲染、机器人、自动驾驶和AR/VR技术至关重要。传统依赖激光雷达等硬件传感器的方法受限于高成本、低分辨率和环境敏感性,难以在真实场景中广泛应用。近年来,基于视觉的方法提供了有前景的替代方案,但因模型容量有限或依赖领域特定的小规模数据集,仍面临泛化性和稳定性挑战。其他领域的缩放定律与基础模型的兴起,推动了“深度基础模型”的发展:在大规模数据上训练的深度神经网络,具备强零样本泛化能力。本文综述了单目、双目、多视角及单目视频设置下深度估计的深度学习架构与范式演进,探讨这些模型解决现有问题的潜力,并系统梳理可促进其发展的大规模数据集。通过识别关键架构与训练策略,本文旨在指明构建稳健深度基础模型的路径,为未来研究与应用提供洞见。
原文摘要 · Abstract (English)
Depth estimation is a fundamental task in 3D computer vision, crucial for applications such as 3D reconstruction, free-viewpoint rendering, robotics, autonomous driving, and AR/VR technologies. Traditional methods relying on hardware sensors like LiDAR are often limited by high costs, low resolution, and environmental sensitivity, limiting their applicability in real-world scenarios. Recent advances in vision-based methods offer a promising alternative, yet they face challenges in generalization and stability due to either the low-capacity model architectures or the reliance on domain-specific and small-scale datasets. The emergence of scaling laws and foundation models in other domains has inspired the development of "depth foundation models": deep neural networks trained on large datasets with strong zero-shot generalization capabilities. This paper surveys the evolution of deep learning architectures and paradigms for depth estimation across the monocular, stereo, multi-view, and monocular video settings. We explore the potential of these models to address existing challenges and provide a comprehensive overview of large-scale datasets that can facilitate their development. By identifying key architectures and training strategies, we aim to highlight the path towards robust depth foundation models, offering insights into their future research and applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。