arXiv:2410.00736cs.RO2024-10被引 2

融合毫米波雷达与单目相机,提升移动机器人在低纹理环境下的深度估计精度。

Radar Meets Vision: Robustifying Monocular Metric Depth Prediction for Mobile Robotics

  • 将低成本雷达数据注入单目深度模型输入层,增强感知鲁棒性。
  • 在未见真实场景中,深度预测的绝对相对误差降低9%-64%。
  • 适用于工业与户外等视觉失效场景,适合移动机器人开发者参考。

移动机器人需精确可靠的深度感知以理解并交互环境。尽管现有传感方式部分满足需求,近期单目深度估计利用单目相机的信息丰富性、低成本与简便性取得显著进展,主要应用于车载与室内场景。然而,机器人常面临尺度线索不足、自相似外观及低纹理环境。本文将低成本毫米波雷达的测量结果编码至先进单目深度估计模型的输入空间。尽管雷达点云极度稀疏,本方法在工业与户外实验中均表现出良好泛化性与鲁棒性。在多个未见过的真实验证数据集上,深度预测的绝对相对误差降低9%-64%。重要的是,所有实验中性能指标保持一致,且在当前纯视觉方法失效的场景与深度范围内表现稳定。此外,针对移动机器人领域训练数据匮乏问题,提出基于摄影测量数据合成逼真渲染数据集的新方法,模拟雷达传感器观测用于训练。代码、数据集与预训练模型已开源:https://github.com/ethz-asl/radarmeetsvision。

原文摘要 · Abstract (English)

Mobile robots require accurate and robust depth measurements to understand and interact with the environment. While existing sensing modalities address this problem to some extent, recent research on monocular depth estimation has leveraged the information richness, yet low cost and simplicity of monocular cameras. These works have shown significant generalization capabilities, mainly in automotive and indoor settings. However, robots often operate in environments with limited scale cues, self-similar appearances, and low texture. In this work, we encode measurements from a low-cost mmWave radar into the input space of a state-of-the-art monocular depth estimation model. Despite the radar's extreme point cloud sparsity, our method demonstrates generalization and robustness across industrial and outdoor experiments. Our approach reduces the absolute relative error of depth predictions by 9-64% across a range of unseen, real-world validation datasets. Importantly, we maintain consistency of all performance metrics across all experiments and scene depths where current vision-only approaches fail. We further address the present deficit of training data in mobile robotics environments by introducing a novel methodology for synthesizing rendered, realistic learning datasets based on photogrammetric data that simulate the radar sensor observations for training. Our code, datasets, and pre-trained networks are made available at https://github.com/ethz-asl/radarmeetsvision.

深度估计多模态融合移动机器人雷达

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。