arXiv:2510.10876cs.CV2025-10

构建合成点云数据集RareBoost3D,解决稀有类别样本不足问题

rareboost3d: a synthetic lidar dataset with enhanced rare classes

  • 构建合成点云数据集,大幅增加稀有类别的样本数量
  • 跨域语义对齐方法使真实数据分割性能提升显著
  • 适合自动驾驶感知模型训练与长尾分布研究

真实世界点云数据集在基于激光雷达的感知技术发展中发挥了重要作用,如自动驾驶中的目标分割。然而,由于某些稀有类别实例数量有限,长尾问题仍是现有数据集的主要挑战。为解决此问题,我们提出一种新型合成点云数据集RareBoost3D,通过提供比真实数据集多得多的稀有类别实例来弥补现有数据集的不足。为进一步有效利用合成与真实数据,我们还提出一种跨域语义对齐方法CSC loss,实现不同域中同类物体特征表示的一致性。实验结果表明,该对齐方法显著提升了基于真实数据的激光雷达点云分割模型性能。

原文摘要 · Abstract (English)

Real-world point cloud datasets have made significant contributions to the development of LiDAR-based perception technologies, such as object segmentation for autonomous driving. However, due to the limited number of instances in some rare classes, the long-tail problem remains a major challenge in existing datasets. To address this issue, we introduce a novel, synthetic point cloud dataset named RareBoost3D, which complements existing real-world datasets by providing significantly more instances for object classes that are rare in real-world datasets. To effectively leverage both synthetic and real-world data, we further propose a cross-domain semantic alignment method named CSC loss that aligns feature representations of the same class across different domains. Experimental results demonstrate that this alignment significantly enhances the performance of LiDAR point cloud segmentation models over real-world data.

点云数据合成数据长尾问题自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。