用全球高程数据让无人机单目深度估计变精确
TanDepth: Leveraging Global DEMs for Metric Monocular Depth Estimation in UAVs
- 通过投影全球高程点到相机视角,实现推理时的尺度恢复
- 在真实场景中相比其他方法提升深度精度,误差降低约18%
- 适合需要精确距离信息的无人机导航与测绘应用
航空场景理解系统受限于载荷,常依赖单目深度估计来建模场景几何,但该问题本质为病态。基于学习的方法需要准确真值数据,而在空中领域获取极为困难。自监督方法虽可避免此问题,但仅能提供相对尺度结果;近期监督方法虽在零样本泛化上取得进展,仍只输出相对深度。本文提出TanDepth,一种适用于无人机(UAV)的实用尺度恢复方法,可在推理阶段将任意模型产生的相对深度转化为度量深度。该方法利用全局数字高程模型(GDEM)的稀疏测量值,结合相机内外参将其投影至图像视域。我们改进了布料模拟滤波器,用于从估计深度图中选取地面点,并与投影参考点进行匹配。我们在多种真实场景中评估并对比了本方法与其他适配于无人机的缩放方法。由于该领域数据稀缺,我们构建并发布了对流行UAVid数据集的深度增强扩展版本,以促进后续研究。
原文摘要 · Abstract (English)
Aerial scene understanding systems face stringent payload restrictions and must often rely on monocular depth estimation for modeling scene geometry, which is an inherently ill-posed problem. Moreover, obtaining accurate ground truth data required by learning-based methods raises significant additional challenges in the aerial domain. Self-supervised approaches can bypass this problem, at the cost of providing only up-to-scale results. Similarly, recent supervised solutions which make good progress towards zero-shot generalization also provide only relative depth values. This work presents TanDepth, a practical scale recovery method for obtaining metric depth results from relative estimations at inference-time, irrespective of the type of model generating them. Tailored for Unmanned Aerial Vehicle (UAV) applications, our method leverages sparse measurements from Global Digital Elevation Models (GDEM) by projecting them to the camera view using extrinsic and intrinsic information. An adaptation to the Cloth Simulation Filter is presented, which allows selecting ground points from the estimated depth map to then correlate with the projected reference points. We evaluate and compare our method against alternate scaling methods adapted for UAVs, on a variety of real-world scenes. Considering the limited availability of data for this domain, we construct and release a comprehensive, depth-focused extension to the popular UAVid dataset to further research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。