用单目相机和惯性传感器实现零样本度量深度估计,支持无人机避障。
Zero-Shot Metric Depth Estimation via Monocular Visual-Inertial Rescaling for Autonomous Aerial Navigation
- 基于视觉惯性系统生成稀疏3D特征图,通过轻量级重缩放策略获得度量深度
- 在模拟环境中验证,最优方案采用单调样条拟合,真实飞行中达15Hz实时输出
- 适合计算资源受限的无人机,无需额外训练即可部署于自主导航场景
本文提出一种从单目RGB图像与惯性测量单元(IMU)推断度量深度的方法。为实现自主飞行中的碰撞规避,现有方法或依赖重型传感器(如激光雷达或双目相机),或需大量数据且针对特定场景微调单目度量深度模型。相比之下,本文提出若干轻量级零样本重缩放策略,利用视觉惯性导航系统构建的稀疏3D特征图,将相对深度估计转换为度量深度。在多种模拟环境对比了不同策略的精度,表现最佳的方法采用单调样条拟合,并在计算资源受限的四旋翼无人机上实现实时部署。系统实现15 Hz的机载度量深度估计,结合基于运动原语的规划器成功完成避障任务。
原文摘要 · Abstract (English)
This paper presents a methodology to predict metric depth from monocular RGB images and an inertial measurement unit (IMU). To enable collision avoidance during autonomous flight, prior works either leverage heavy sensors (e.g., LiDARs or stereo cameras) or data-intensive and domain-specific fine-tuning of monocular metric depth estimation methods. In contrast, we propose several lightweight zero-shot rescaling strategies to obtain metric depth from relative depth estimates via the sparse 3D feature map created using a visual-inertial navigation system. These strategies are compared for their accuracy in diverse simulation environments. The best performing approach, which leverages monotonic spline fitting, is deployed in the real-world on a compute-constrained quadrotor. We obtain on-board metric depth estimates at 15 Hz and demonstrate successful collision avoidance after integrating the proposed method with a motion primitives-based planner.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。