用2D模型的开放词汇能力提升3D场景分割,无需重新训练。
SAS: Segment Any 3D Scene with Integrated 2D Priors
- 通过文本对齐将多个2D模型映射到统一空间。
- 用扩散模型量化2D模型识别不同类别的能力。
- 在多个数据集上超越现有方法,适合开放词汇3D分割任务。
3D模型的开放词汇能力日益重要,传统方法因训练类别固定,难以识别复杂动态3D场景中的未见物体。本文提出SAS,一种简单有效的方案,将多个2D模型的开放词汇能力迁移到3D领域。首先,通过文本作为桥梁,实现不同2D模型间的嵌入空间对齐;其次,利用扩散模型显式量化2D模型对各类别的识别能力;随后,基于构建的模型能力融合来自不同2D模型的点云特征;最后,通过特征蒸馏将整合后的2D开放词汇能力迁移至3D域。SAS在ScanNet v2、Matterport3D和nuScenes等多个数据集上显著优于以往方法,其泛化能力在下游任务(如高斯分割和实例分割)中也得到验证。
原文摘要 · Abstract (English)
The open vocabulary capability of 3D models is increasingly valued, as traditional methods with models trained with fixed categories fail to recognize unseen objects in complex dynamic 3D scenes. In this paper, we propose a simple yet effective approach, SAS, to integrate the open vocabulary capability of multiple 2D models and migrate it to 3D domain. Specifically, we first propose Model Alignment via Text to map different 2D models into the same embedding space using text as a bridge. Then, we propose Annotation-Free Model Capability Construction to explicitly quantify the 2D model's capability of recognizing different categories using diffusion models. Following this, point cloud features from different 2D models are fused with the guide of constructed model capabilities. Finally, the integrated 2D open vocabulary capability is transferred to 3D domain through feature distillation. SAS outperforms previous methods by a large margin across multiple datasets, including ScanNet v2, Matterport3D, and nuScenes, while its generalizability is further validated on downstream tasks, e.g., gaussian segmentation and instance segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。