arXiv:2508.00580cs.ROcs.AI2025-08被引 1

融合多模态图像实现火星车复杂地形语义分割

OmniUnet: A Multimodal Network for Unstructured Terrain Segmentation on Planetary Rovers Using RGB, Depth, and Thermal Imagery

  • 基于Transformer设计多模态网络,整合RGB、深度与热成像数据
  • 在真实沙漠地形上达到80.37%像素准确率,推理时间仅673毫秒
  • 开源数据集与代码,助力行星机器人感知研究

机器人在非结构化环境中导航需依赖多模态感知系统以保障安全。多模态信息可融合不同传感器的互补特征,但需专门设计的机器学习算法处理异构数据。火星探测中,热成像因土壤类型间热特性差异而对评估地形安全性具有价值。本文提出OmniUnet,一种基于Transformer的神经网络架构,用于结合RGB、深度与热成像(RGB-D-T)进行语义分割。通过3D打印定制多模态传感器支架,安装于火星车测试平台MaRTA,在西班牙北部巴登纳斯半荒漠区采集真实地形数据,该区域模拟火星表面,包含沙地、基岩和压实土壤。部分数据经人工标注,用于监督训练。模型在定量与定性评估中表现良好,像素准确率达80.37%,在复杂未结构化地形分割中效果显著。在资源受限设备Jetson Orin Nano上平均推理时间为673毫秒,证实其适合车载部署。网络软件与标注数据集已公开,支持未来行星机器人多模态感知研究。

原文摘要 · Abstract (English)

Robot navigation in unstructured environments requires multimodal perception systems that can support safe navigation. Multimodality enables the integration of complementary information collected by different sensors. However, this information must be processed by machine learning algorithms specifically designed to leverage heterogeneous data. Furthermore, it is necessary to identify which sensor modalities are most informative for navigation in the target environment. In Martian exploration, thermal imagery has proven valuable for assessing terrain safety due to differences in thermal behaviour between soil types. This work presents OmniUnet, a transformer-based neural network architecture for semantic segmentation using RGB, depth, and thermal (RGB-D-T) imagery. A custom multimodal sensor housing was developed using 3D printing and mounted on the Martian Rover Testbed for Autonomy (MaRTA) to collect a multimodal dataset in the Bardenas semi-desert in northern Spain. This location serves as a representative environment of the Martian surface, featuring terrain types such as sand, bedrock, and compact soil. A subset of this dataset was manually labeled to support supervised training of the network. The model was evaluated both quantitatively and qualitatively, achieving a pixel accuracy of 80.37% and demonstrating strong performance in segmenting complex unstructured terrain. Inference tests yielded an average prediction time of 673 ms on a resource-constrained computer (Jetson Orin Nano), confirming its suitability for on-robot deployment. The software implementation of the network and the labeled dataset have been made publicly available to support future research in multimodal terrain perception for planetary robotics.

多模态感知地形分割火星探测嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。