提出果园场景单目深度估计新方法,显著提升精度。
OrchardDepth: Precise Metric Depth Estimation of Orchard Scene from Monocular Camera Images
- 设计新训练策略,融合稠密深度图与稀疏点的一致性正则化
- 在果园场景中将深度估计RMSE从1.5337降至0.6738
- 专为果园/葡萄园环境优化,适合农业机器人应用
单目深度估计是机器人感知的基础任务。近年来,随着神经网络模型和数据集的发展,该任务的性能和效率显著提升。然而,多数研究集中于城市等密集场景,现有户外基准数据集多面向自动驾驶,与果园/葡萄园等农业场景差异巨大,难以支持农业领域研究。为此,本文提出OrchardDepth,填补了单目相机在果园/葡萄园环境中进行精确度量深度估计的空白。同时,提出一种新的重训练方法,通过监控稠密深度图与稀疏点之间的一致性正则化来提升训练效果。实验表明,该方法将果园场景下的深度估计RMSE从1.5337降低至0.6738,验证了其有效性。
原文摘要 · Abstract (English)
Monocular depth estimation is a rudimentary task in robotic perception. Recently, with the development of more accurate and robust neural network models and different types of datasets, monocular depth estimation has significantly improved performance and efficiency. However, most of the research in this area focuses on very concentrated domains. In particular, most of the benchmarks in outdoor scenarios belong to urban environments for the improvement of autonomous driving devices, and these benchmarks have a massive disparity with the orchard/vineyard environment, which is hardly helpful for research in the primary industry. Therefore, we propose OrchardDepth, which fills the gap in the estimation of the metric depth of the monocular camera in the orchard/vineyard environment. In addition, we present a new retraining method to improve the training result by monitoring the consistent regularization between dense depth maps and sparse points. Our method improves the RMSE of depth estimation in the orchard environment from 1.5337 to 0.6738, proving our method's validation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。