多模态自监督学习提升机器人在复杂地形的通行能力评估
Scene-Agnostic Traversability Labeling and Estimation via a Multimodal Self-supervised Framework
- 融合足印、激光雷达和相机数据生成通行性标签
- 自动标注达到约88%交并比,优于现有自监督方法1.6-3.5%
- 适合需要鲁棒地形理解的移动机器人研发人员
通行性估计对机器人在多样化地形与环境中的导航至关重要。尽管近期自监督学习方法取得良好效果,但往往无法捕捉不可通行区域的特征。此外,多数工作仅依赖单一模态,忽视了异构传感模态的互补优势。为此,我们提出一种多模态自监督框架,用于通行性标注与估计。首先,通过融合足印、激光雷达和相机数据作为提示,利用视觉基础模型生成同时包含语义与几何线索的通行性标签。随后,基于这些标签训练一个双流网络,以解耦方式联合学习不同模态信息,增强对多样通行模式的识别能力。此外,引入稀疏激光雷达监督以缓解伪标签带来的噪声。在城市、非铺装道路及校园环境的大量实验表明,所提自动标注方法在多个数据集上均达到约88% IoU。相比现有自监督最先进方法,我们的多模态通行性估计网络在所有测试数据集上均提升1.6%-3.5% IoU。
原文摘要 · Abstract (English)
Traversability estimation is critical for enabling robots to navigate across diverse terrains and environments. While recent self-supervised learning methods achieve promising results, they often fail to capture the characteristics of non-traversable regions. Moreover, most prior works concentrate on a single modality, overlooking the complementary strengths offered by integrating heterogeneous sensory modalities for more robust traversability estimation. To address these limitations, we propose a multimodal self-supervised framework for traversability labeling and estimation. First, our annotation pipeline integrates footprint, LiDAR, and camera data as prompts for a vision foundation model, generating traversability labels that account for both semantic and geometric cues. Then, leveraging these labels, we train a dual-stream network that jointly learns from different modalities in a decoupled manner, enhancing its capacity to recognize diverse traversability patterns. In addition, we incorporate sparse LiDAR-based supervision to mitigate the noise introduced by pseudo labels. Finally, extensive experiments conducted across urban, off-road, and campus environments demonstrate the effectiveness of our approach. The proposed automatic labeling method consistently achieves around 88% IoU across diverse datasets. Compared to existing self-supervised state-of-the-art methods, our multimodal traversability estimation network yields consistently higher IoU, improving by 1.6-3.5% on all evaluated datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。