用图像数据生成激光雷达的伪标签,无需人工标注即可提升360°感知性能。
XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation
- 通过相机检测构建语义参数地图,实现跨模态知识迁移。
- 在3D语义分割上提升32.3 mIoU,2D分割提升19.5 mIoU。
- 适用于无重叠传感器场景,适合自动驾驶多模态训练数据生成。
在开放世界基础模型尚未达到专用方法性能之前,深度学习系统仍依赖特定任务与传感器的数据。为弥合可用数据集与部署域之间的差距,领域自适应策略被广泛采用。本文提出XD-MAP,一种将图像数据集中的传感器特异性知识迁移到激光雷达(LiDAR)这一完全不同感知域的新方法。该方法利用相机图像上的检测结果构建语义参数地图,地图元素可生成目标域的伪标签,无需任何人工标注。与以往方法不同,本方法不要求传感器间直接重叠,且能将前视相机的感知范围扩展至360°全景。在大规模道路特征数据集上,XD-MAP相较于单次基准方法,在2D语义分割上提升19.5 mIoU,2D全景分割提升19.5 PQth,3D语义分割提升32.3 mIoU。结果表明,该方法可在无需人工标注的情况下,在激光雷达数据上实现优异性能。
原文摘要 · Abstract (English)
Until open-world foundation models match the performance of specialized approaches, deep learning systems remain dependent on task- and sensor-specific data availability. To bridge the gap between available datasets and deployment domains, domain adaptation strategies are widely used. In this work, we propose XD-MAP, a novel approach to transfer sensor-specific knowledge from an image dataset to LiDAR, an entirely different sensing domain. Our method leverages detections on camera images to create a semantic parametric map. The map elements are modeled to produce pseudo labels in the target domain without any manual annotation effort. Unlike previous domain transfer approaches, our method does not require direct overlap between sensors and enables extending the angular perception range from a front-view camera to a full 360° view. On our large-scale road feature dataset, XD-MAP outperforms single shot baseline approaches by +19.5 mIoU for 2D semantic segmentation, +19.5 PQth for 2D panoptic segmentation, and +32.3 mIoU in 3D semantic segmentation. The results demonstrate the effectiveness of our approach achieving strong performance on LiDAR data without any manual labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。