通过融合数据增强与点云模型,提升自动驾驶3D语义分割精度
The 2nd Place Solution from the 3D Semantic Segmentation Track in the 2024 Waymo Open Dataset Challenge
- 采用MinkUNet结合LaserMix与PolarMix进行场景级数据增强
- 在Waymo数据集上实现2024挑战赛第二名的分割性能
- 适合关注自动驾驶感知与点云数据增强的研究者
3D语义分割是驾驶感知中至关重要的任务,学习型模型对密集3D环境的准确感知常决定自动驾驶车辆的安全运行。然而,现有基于激光雷达的3D语义分割数据集由连续采集的激光雷达扫描构成,存在长尾分布且训练多样性不足的问题。本文介绍MixSeg3D,一种强点云分割模型与先进3D数据混合策略的复杂组合。具体而言,该方法将MinkUNet系列模型与LaserMix和PolarMix两种场景级数据增强方法结合,沿本体场景的俯仰角与方位角方向混合激光雷达点云。通过实证实验,我们证明了MixSeg3D优于基线及已有方法。本团队在2024年Waymo Open Dataset挑战赛的3D语义分割赛道中获得第二名。
原文摘要 · Abstract (English)
3D semantic segmentation is one of the most crucial tasks in driving perception. The ability of a learning-based model to accurately perceive dense 3D surroundings often ensures the safe operation of autonomous vehicles. However, existing LiDAR-based 3D semantic segmentation databases consist of sequentially acquired LiDAR scans that are long-tailed and lack training diversity. In this report, we introduce MixSeg3D, a sophisticated combination of the strong point cloud segmentation model with advanced 3D data mixing strategies. Specifically, our approach integrates the MinkUNet family with LaserMix and PolarMix, two scene-scale data augmentation methods that blend LiDAR point clouds along the ego-scene's inclination and azimuth directions. Through empirical experiments, we demonstrate the superiority of MixSeg3D over the baseline and prior arts. Our team achieved 2nd place in the 3D semantic segmentation track of the 2024 Waymo Open Dataset Challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。