arXiv:2605.08293cs.CV2026-05

无需标注数据,通过多级蒸馏与图扩散实现自动驾驶3D语义分割。

Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion

论文配图:Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion
图 1 · 摘自论文原文
  • 分层掩码级联增强小物体在不同尺度下的保留能力。
  • 多级蒸馏提升区域一致性与跨模态区分性,精度提升超9%。
  • 基于重启的图扩散高效传播上下文信息,适合真实驾驶场景。

基于激光雷达的语义分割对自动驾驶感知至关重要,但密集点级标注成本高昂,且长尾户外场景中,小而关键物体在无监督条件下难以发现。现有无监督方法面临三大挑战:在大尺度变化下难以保留小而稀疏物体;跨模态迁移中难以保证区域内一致性和区域间区分性;缺乏高效的特征保持机制用于超点图上的上下文传播。为此,我们提出DDS框架。首先,粗到细的多粒度掩码级联为不同尺度物体提供互补的3D区域线索,改善小物体和稀疏观测物体的保留。其次,区域引导的多级蒸馏通过点级对齐、掩码级原型对齐及原型级对比学习,提升区域一致性与区分性。第三,基于重启的图扩散在超点间高效传播上下文信息,同时锚定优化表示于初始蒸馏特征,避免显式图特征分解。在真实驾驶数据集上的实验表明,DDS优于代表性无监督基线,在oAcc、mAcc和mIoU上分别提升最高2.9%、9.7%和4.1%。结果验证了其在自动驾驶场景中无监督3D场景理解的有效性与可迁移性。

原文摘要 · Abstract (English)

LiDAR-based semantic segmentation is essential for autonomous-driving perception, yet dense point-wise annotations are costly, and long-tailed outdoor scenes make small safety-critical objects difficult to discover without supervision. Existing unsupervised methods face three key challenges: they struggle to preserve small and sparsely observed objects under substantial scale variation, have difficulty enforcing intra-region consistency and inter-region discrimination during cross-modal transfer, and lack an efficient feature-preserving mechanism for contextual propagation over superpoint graphs. We therefore propose DDS, an unsupervised 3D semantic segmentation framework. First, a coarse-to-fine multi-granularity mask cascade provides complementary 3D region cues for objects across different scales, improving the preservation of small and sparsely observed objects. Second, region-guided multi-level distillation transfers self-supervised visual knowledge through point-level alignment, mask-level prototype alignment, and prototype-level contrastive learning, enhancing intra-region consistency and inter-region discrimination. Third, restart-based graph diffusion efficiently propagates contextual information among superpoints while anchoring the refined representation to the initial distilled features and avoiding explicit graph eigendecomposition. Experiments on real-world driving datasets show that DDS outperforms representative unsupervised baselines, improving oAcc, mAcc, and mIoU by up to 2.9%, 9.7%, and 4.1%, respectively. These results demonstrate the effectiveness and transferability of DDS for unsupervised 3D scene understanding in autonomous-driving scenarios.

3D分割无监督学习自动驾驶图扩散

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。