通过点云与实体的对比学习,实现开放词汇3D语义分割的新方法。
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding

- 利用点云视图间实体一致性与语言对齐建模特征
- 在ScanNet上达到当前最佳零样本分割效果
- 适合需要跨场景理解的机器人感知任务
开放词汇3D场景理解对提升智能体的物理认知能力至关重要,使其能够动态解析并交互于真实环境。本文提出MPEC——一种新型的掩码点-实体对比学习方法,用于开放词汇3D语义分割。该方法结合3D实体与语言的对齐关系,以及不同点云视角下的点-实体一致性,以构建具有实体特异性的特征表示。所提方法增强了语义区分能力,提升了实例间的可辨识度,在ScanNet数据集上实现了开放词汇3D语义分割的最先进性能,并展现出优异的零样本场景理解能力。在8个涵盖从低层感知到高层推理任务的数据集上进行的广泛微调实验,验证了所学3D特征的泛化潜力,推动了多种3D场景理解任务的一致性能提升。
原文摘要 · Abstract (English)
Open-vocabulary 3D scene understanding is pivotal for enhancing physical intelligence, as it enables embodied agents to interpret and interact dynamically within real-world environments. This paper introduces MPEC, a novel Masked Point-Entity Contrastive learning method for open-vocabulary 3D semantic segmentation that leverages both 3D entity-language alignment and point-entity consistency across different point cloud views to foster entity-specific feature representations. Our method improves semantic discrimination and enhances the differentiation of unique instances, achieving state-of-the-art results on ScanNet for open-vocabulary 3D semantic segmentation and demonstrating superior zero-shot scene understanding capabilities. Extensive fine-tuning experiments on 8 datasets, spanning from low-level perception to high-level reasoning tasks, showcase the potential of learned 3D features, driving consistent performance gains across varied 3D scene understanding tasks. Project website: https://mpec-3d.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。