为火星车设计实时+离线视觉系统,提升自主导航能力
Deep Learning Aided Vision System for Planetary Rovers
- 实时模块融合增强立体图像与轻量检测模型
- 1-10米内深度误差仅2.26厘米,精度高
- 适合资源受限的深空探测任务使用
本研究提出一种面向行星车的视觉系统,结合实时感知与离线地形重建。实时模块整合CLAHE增强的立体图像、基于YOLOv11n的物体检测和神经网络估计物体距离;离线模块利用Depth Anything V2单目深度估计模型生成深度图,并通过Open3D融合为稠密点云。真实世界距离估计为定性重构提供了可靠的度量上下文。在Chandrayaan 3 NavCam立体图像上的评估显示,神经网络在1至10米范围内达到2.26厘米的中位深度误差;物体检测模型在灰度月面场景中保持精度与召回率的平衡。该架构为自主行星探测提供了可扩展、计算高效的视觉解决方案。
原文摘要 · Abstract (English)
This study presents a vision system for planetary rovers, combining real-time perception with offline terrain reconstruction. The real-time module integrates CLAHE enhanced stereo imagery, YOLOv11n based object detection, and a neural network to estimate object distances. The offline module uses the Depth Anything V2 metric monocular depth estimation model to generate depth maps from captured images, which are fused into dense point clouds using Open3D. Real world distance estimates from the real time pipeline provide reliable metric context alongside the qualitative reconstructions. Evaluation on Chandrayaan 3 NavCam stereo imagery, benchmarked against a CAHV based utility, shows that the neural network achieves a median depth error of 2.26 cm within a 1 to 10 meter range. The object detection model maintains a balanced precision recall tradeoff on grayscale lunar scenes. This architecture offers a scalable, compute-efficient vision solution for autonomous planetary exploration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。