针对自动驾驶自监督预训练中的类别不平衡与场景复杂问题,提出结构化聚类方法提升表征能力。
S3PT: Scene Semantics and Structure Guided Clustering to Boost Self-Supervised Pre-Training for Autonomous Driving
- 融合语义分布与空间结构一致性聚类,增强稀有类别表征
- 在nuScenes等数据集上,分割与3D检测性能显著提升
- 适合自动驾驶领域需鲁棒几何感知的自监督学习任务
近期基于聚类的自监督预训练方法(如DINO和Cribo)在下游检测与分割任务中表现优异。然而,在自动驾驶等真实场景中,存在类别与尺寸分布不均、场景几何复杂等问题。本文提出S3PT——一种基于场景语义与结构引导的聚类方法,以提供更一致的自监督目标。首先,引入语义分布一致聚类,改善摩托车、动物等稀有类别的表征;其次,设计对象多样性一致的空间聚类,应对从大范围背景到行人、交通标志等小物体的尺寸差异;第三,提出深度引导的空间聚类,利用场景几何信息正则化特征学习,优化区域分离。所学表征在nuScenes、nuImages和Cityscapes数据集上的下游语义分割与3D目标检测任务中均有显著提升,并展现出良好的域迁移潜力。
原文摘要 · Abstract (English)
Recent self-supervised clustering-based pre-training techniques like DINO and Cribo have shown impressive results for downstream detection and segmentation tasks. However, real-world applications such as autonomous driving face challenges with imbalanced object class and size distributions and complex scene geometries. In this paper, we propose S3PT a novel scene semantics and structure guided clustering to provide more scene-consistent objectives for self-supervised training. Specifically, our contributions are threefold: First, we incorporate semantic distribution consistent clustering to encourage better representation of rare classes such as motorcycles or animals. Second, we introduce object diversity consistent spatial clustering, to handle imbalanced and diverse object sizes, ranging from large background areas to small objects such as pedestrians and traffic signs. Third, we propose a depth-guided spatial clustering to regularize learning based on geometric information of the scene, thus further refining region separation on the feature level. Our learned representations significantly improve performance in downstream semantic segmentation and 3D object detection tasks on the nuScenes, nuImages, and Cityscapes datasets and show promising domain translation properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。