arXiv:2509.11082cs.CVcs.RO2025-09被引 1

用自监督IMU标签融合图像与激光雷达,生成火星车通行成本图。

Mars Traversability Prediction: A Multi-modal Self-supervised Approach for Costmap Generation

  • 基于DINOv3和FiLM的多模态融合,从视觉与激光数据生成鸟瞰成本图。
  • 在稀疏激光雷达下MAE仅从0.0775升至0.0915,模型鲁棒性强。
  • 适合火星探测任务中高保真仿真与低成本标注场景。

本文提出一种鲁棒的多模态框架,用于预测行星探测车的通行成本图。模型融合摄像头与激光雷达数据,生成鸟瞰视角(BEV)地形成本图,并利用惯性测量单元(IMU)导出的标签进行自监督训练。关键改进包括基于DINOv3的图像编码器、基于FiLM的传感器融合机制,以及结合Huber损失与平滑项的优化目标。实验消融测试(去除图像颜色、遮挡输入、添加噪声)显示,平均绝对误差(MAE)/均方误差(MSE)仅小幅上升(例如,当激光雷达数据稀疏时MAE由约0.0775增至0.0915),表明几何信息主导了学习到的成本分布,模型具有高度鲁棒性。性能差异较小的原因在于IMU标签主要反映地形几何而非语义,且数据多样性有限。本文强调核心贡献:(1)高保真、可复现的仿真环境;(2)基于IMU的自监督标注流程;(3)强大的多模态BEV成本图预测模型。同时讨论了领域泛化与数据集扩展等未来方向。

原文摘要 · Abstract (English)

We present a robust multi-modal framework for predicting traversability costmaps for planetary rovers. Our model fuses camera and LiDAR data to produce a bird's-eye-view (BEV) terrain costmap, trained self-supervised using IMU-derived labels. Key updates include a DINOv3-based image encoder, FiLM-based sensor fusion, and an optimization loss combining Huber and smoothness terms. Experimental ablations (removing image color, occluding inputs, adding noise) show only minor changes in MAE/MSE (e.g. MAE increases from ~0.0775 to 0.0915 when LiDAR is sparsified), indicating that geometry dominates the learned cost and the model is highly robust. We attribute the small performance differences to the IMU labeling primarily reflecting terrain geometry rather than semantics and to limited data diversity. Unlike prior work claiming large gains, we emphasize our contributions: (1) a high-fidelity, reproducible simulation environment; (2) a self-supervised IMU-based labeling pipeline; and (3) a strong multi-modal BEV costmap prediction model. We discuss limitations and future work such as domain generalization and dataset expansion.

多模态自监督成本图火星探测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。