arXiv:2607.16366cs.ROcs.AI2026-07

融合可见光、深度与热成像,提升火星车在复杂地形的导航可靠性。

PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments

论文配图:PRISM: Multimodal Terrain Mapping for Rover Navigation in Unstructured Environments
图 1 · 摘自论文原文
  • 用RGB-D-T多模态数据+OmniUnet网络进行地表语义分割。
  • 在两个新标注数据集上验证,实测可在嵌入式设备高效运行。
  • 适合资源受限的机器人导航系统,尤其适用于外星探测任务。

在非结构化环境中,机器人导航需要强大的态势感知能力以安全穿越陡坡、岩石等地形障碍。为此,感知系统越来越多依赖多模态传感器融合。将热成像与标准的可见光和深度传感器结合,可显著提升地形区分能力,从而增强地图算法的可靠性。本文提出PRISM多模态感知系统,用于非结构化环境下的地形建图。该系统采用定制传感器套件,同步采集对齐的RGB、深度与热成像(RGB-D-T)数据。核心为OmniUnet——一种专为多模态语义地形分割设计的视觉变压器网络。通过两个新标注数据集(BASEPROD和LAENTIEC)验证,并在实地实验中展示其实际应用价值。部署于资源受限的嵌入式计算机,PRISM能高效生成可直接驱动机器人导航控制(GNC)子系统的可通行性地图。

原文摘要 · Abstract (English)

Robotic navigation in unstructured environments requires robust situational awareness to safely traverse hazards such as steep slopes and rocky terrain. To address this challenge, perception systems increasingly rely on multimodal sensor fusion. Specifically, integrating thermal imagery with standard optical and depth sensors enhances terrain differentiation, directly improving the reliability of mapping algorithms. This paper presents PRISM, a multimodal perception system for terrain mapping in unstructured settings. PRISM leverages a custom sensor suite to capture aligned RGB, depth, and thermal (RGB-D-T) imagery. At its core is OmniUnet, a novel vision transformer-based network specifically designed for multimodal semantic terrain segmentation. We validated the proposed system using two newly annotated datasets (BASEPROD and LAENTIEC) and demonstrate its real-world applicability through physical field experiments. Deployed on a resource-constrained embedded computer, PRISM efficiently generates traversability maps that directly enable autonomous navigation via a rover's Guidance, Navigation, and Control (GNC) subsystem.

多模态感知地形建图机器人导航嵌入式部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。