arXiv:2409.00845cs.CV2024-09ECCV被引 7

通过关系蒸馏对齐2D与3D表征结构,提升自动驾驶点云分割性能。

Image-to-Lidar Relational Distillation for Autonomous Driving Data

论文配图:Image-to-Lidar Relational Distillation for Autonomous Driving Data
图 1 · 摘自论文原文
  • 设计跨模态与同模内约束的关系蒸馏框架
  • 零样本分割性能显著优于对比蒸馏方法
  • 适合研究自动驾驶3D感知与多模态迁移学习的学者

在大规模多样化的多模态数据集上预训练的2D基础模型,在极少或无需下游监督的情况下即可出色完成2D任务,得益于其强大的表征能力。2D到3D的蒸馏框架已将这些能力扩展至3D模型。然而,针对自动驾驶数据集的3D表征蒸馏面临自相似性、类别不平衡和点云稀疏性等挑战,导致对比蒸馏在零样本学习中效果受限。尽管基于相似性的方法能提升零样本性能,但通常产生判别性较弱的表征,损害了少样本性能。本文研究了先进蒸馏框架下2D与3D表征间的结构差异,揭示二者存在显著不匹配,且该结构差距与零样本和少样本3D语义分割的蒸馏效果呈负相关。为弥合此差距,提出一种施加同模与跨模约束的关系蒸馏框架,使蒸馏出的3D表征更贴近2D表征结构。该对齐显著提升了零样本分割表现。此外,所提关系损失在分布内与分布外的少样本分割任务中均持续改善3D表征质量,优于依赖相似性损失的方法。

原文摘要 · Abstract (English)

Pre-trained on extensive and diverse multi-modal datasets, 2D foundation models excel at addressing 2D tasks with little or no downstream supervision, owing to their robust representations. The emergence of 2D-to-3D distillation frameworks has extended these capabilities to 3D models. However, distilling 3D representations for autonomous driving datasets presents challenges like self-similarity, class imbalance, and point cloud sparsity, hindering the effectiveness of contrastive distillation, especially in zero-shot learning contexts. Whereas other methodologies, such as similarity-based distillation, enhance zero-shot performance, they tend to yield less discriminative representations, diminishing few-shot performance. We investigate the gap in structure between the 2D and the 3D representations that result from state-of-the-art distillation frameworks and reveal a significant mismatch between the two. Additionally, we demonstrate that the observed structural gap is negatively correlated with the efficacy of the distilled representations on zero-shot and few-shot 3D semantic segmentation. To bridge this gap, we propose a relational distillation framework enforcing intra-modal and cross-modal constraints, resulting in distilled 3D representations that closely capture the structure of the 2D representation. This alignment significantly enhances 3D representation performance over those learned through contrastive distillation in zero-shot segmentation tasks. Furthermore, our relational loss consistently improves the quality of 3D representations in both in-distribution and out-of-distribution few-shot segmentation tasks, outperforming approaches that rely on the similarity loss.

自动驾驶点云分割多模态蒸馏关系学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。