针对月面小物体检测难题,提出高效多模态融合模型SCAFusion。
SCAFusion: A Multimodal 3D Detection Framework for Small Object Detection in Lunar Surface Exploration
- 引入认知适配器与对比对齐模块,提升视觉与激光雷达特征一致性。
- 在模拟月面环境实现90.93% mAP,小目标检测性能提升11.5%。
- 专为月球探测设计,适合深空机器人自主导航场景。
可靠且精确地检测小而不规则的物体(如陨石碎片和岩石)对于月面自主导航与操作至关重要。现有面向地球自动驾驶的多模态3D感知方法在非地球环境中表现不佳,主要由于特征对齐差、多模态协同弱以及小目标检测能力不足。本文提出SCAFusion,一种专为月球机器人任务设计的多模态3D目标检测模型。基于BEVFusion框架,SCAFusion引入认知适配器以高效微调相机主干网络,设计对比对齐模块增强相机与激光雷达特征一致性,加入相机辅助训练分支强化视觉表征,并提出专为小不规则目标优化的分段坐标注意力机制。模型参数与计算开销几乎无增加,在nuScenes验证集上达到69.7% mAP和72.1% NDS,分别比基线提升5.0%和2.7%。在基于Isaac Sim构建的模拟月面环境中,mAP达90.93%,优于基线11.5%,尤其在检测类似陨石的小障碍物方面表现突出。
原文摘要 · Abstract (English)
Reliable and precise detection of small and irregular objects, such as meteor fragments and rocks, is critical for autonomous navigation and operation in lunar surface exploration. Existing multimodal 3D perception methods designed for terrestrial autonomous driving often underperform in off world environments due to poor feature alignment, limited multimodal synergy, and weak small object detection. This paper presents SCAFusion, a multimodal 3D object detection model tailored for lunar robotic missions. Built upon the BEVFusion framework, SCAFusion integrates a Cognitive Adapter for efficient camera backbone tuning, a Contrastive Alignment Module to enhance camera LiDAR feature consistency, a Camera Auxiliary Training Branch to strengthen visual representation, and most importantly, a Section aware Coordinate Attention mechanism explicitly designed to boost the detection performance of small, irregular targets. With negligible increase in parameters and computation, our model achieves 69.7% mAP and 72.1% NDS on the nuScenes validation set, improving the baseline by 5.0% and 2.7%, respectively. In simulated lunar environments built on Isaac Sim, SCAFusion achieves 90.93% mAP, outperforming the baseline by 11.5%, with notable gains in detecting small meteor like obstacles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。