arXiv:2501.19035cs.CV2025-01被引 9

用改进版CARLA生成了与真实数据对齐的合成激光雷达语义分割数据集。

SynthmanticLiDAR: A Synthetic Dataset for Semantic Segmentation on LiDAR Imaging

  • 基于CARLA改装模拟器,支持自定义类别与分布
  • 合成数据可提升多种算法性能,优于纯真实数据训练
  • 适合自动驾驶感知研究者快速获取标注数据

激光雷达语义分割在感知系统和自动驾驶中日益重要,但真实数据采集与标注成本高昂。尽管已有如SemanticKITTI等人工标注数据集,仿真工具(如CARLA)可按需生成合成数据。本文针对语义分割需求改进了CARLA模拟器,新增类别、统一对象标签与SemanticKITTI对齐,并支持调整类别分布。基于此,我们构建了SynthmanticLiDAR合成数据集,其设计目标与SemanticKITTI一致。通过简单迁移学习验证,将该数据集加入训练可显著提升多个语义分割算法的整体性能,证明其有效性。相关数据集与模拟器已开源:https://github.com/vpulab/SynthmanticLiDAR。

原文摘要 · Abstract (English)

Semantic segmentation on LiDAR imaging is increasingly gaining attention, as it can provide useful knowledge for perception systems and potential for autonomous driving. However, collecting and labeling real LiDAR data is an expensive and time-consuming task. While datasets such as SemanticKITTI have been manually collected and labeled, the introduction of simulation tools such as CARLA, has enabled the creation of synthetic datasets on demand. In this work, we present a modified CARLA simulator designed with LiDAR semantic segmentation in mind, with new classes, more consistent object labeling with their counterparts from real datasets such as SemanticKITTI, and the possibility to adjust the object class distribution. Using this tool, we have generated SynthmanticLiDAR, a synthetic dataset for semantic segmentation on LiDAR imaging, designed to be similar to SemanticKITTI, and we evaluate its contribution to the training process of different semantic segmentation algorithms by using a naive transfer learning approach. Our results show that incorporating SynthmanticLiDAR into the training process improves the overall performance of tested algorithms, proving the usefulness of our dataset, and therefore, our adapted CARLA simulator. The dataset and simulator are available in https://github.com/vpulab/SynthmanticLiDAR.

激光雷达语义分割合成数据自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。