用少量标注数据实现越野环境下的鲁棒3D语义地图构建
Few-shot Semantic Learning for Robust Multi-Biome 3D Semantic Mapping in Off-Road Environments
- 基于预训练ViT,仅需<500张图像即可完成少样本语义分割
- 在Yamaha和Rellis数据集上达到66.6和67.2的mIoU,零样本跨生物群落泛化性能达52.9和55.5
- 创新范围融合机制支持动态障碍物与复杂地形的3D语义建图,适合高阶自动驾驶系统
越野环境因无序地形、传感条件退化及生物群落间域偏移,给高速自主导航带来显著感知挑战。当需要大量真实标注数据时,跨环境与生物群落的语义学习尤为困难。本文提出一种方法:利用预训练视觉变换器(ViT),在小规模(<500张图像)、稀疏且粗略标注(<30%像素)的多生物群落数据集上进行微调,实现2D语义分割。这些分割结果通过新颖的基于距离的度量融合,并聚合为3D语义体素地图。我们在Yamaha(52.9 mIoU)和Rellis(55.5 mIoU)数据集上展示了零样本跨生物群落2D语义分割能力;在现有数据上采用少样本粗略标注后,Yamaha(66.6 mIoU)和Rellis(67.2 mIoU)性能进一步提升。此外,我们验证了基于距离的语义融合方法在处理常见越野障碍如突发障碍、悬垂物和水体特征方面的可行性。
原文摘要 · Abstract (English)
Off-road environments pose significant perception challenges for high-speed autonomous navigation due to unstructured terrain, degraded sensing conditions, and domain-shifts among biomes. Learning semantic information across these conditions and biomes can be challenging when a large amount of ground truth data is required. In this work, we propose an approach that leverages a pre-trained Vision Transformer (ViT) with fine-tuning on a small (<500 images), sparse and coarsely labeled (<30% pixels) multi-biome dataset to predict 2D semantic segmentation classes. These classes are fused over time via a novel range-based metric and aggregated into a 3D semantic voxel map. We demonstrate zero-shot out-of-biome 2D semantic segmentation on the Yamaha (52.9 mIoU) and Rellis (55.5 mIoU) datasets along with few-shot coarse sparse labeling with existing data for improved segmentation performance on Yamaha (66.6 mIoU) and Rellis (67.2 mIoU). We further illustrate the feasibility of using a voxel map with a range-based semantic fusion approach to handle common off-road hazards like pop-up hazards, overhangs, and water features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。