arXiv:2502.02548cs.CV2025-02CVPR被引 22

构建超大规模3D语义分割数据集与模型,支持开放词汇理解。

Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation

  • 自动生成高质量3D掩码-文本配对数据,融合视觉语言模型。
  • 创建560万对标注的Mosaic3D数据集,规模超现有数据数倍。
  • 新模型在多个基准上达最优,适合开放词汇3D理解任务。

我们通过引入新型数据生成流程与训练框架,解决开放词汇3D场景理解问题。方法满足三个关键训练需求:精确的3D区域分割、全面的文本描述和充足的数据库规模。利用先进的开放词汇图像分割模型与区域感知视觉-语言模型,开发出自动化流程,生成高质量3D掩码-文本对。该流程应用于多个3D场景数据集,构建了包含超过3万幅标注场景、560万对掩码-文本对的Mosaic3D-5.6M数据集,显著大于现有数据集。基于此数据,提出Mosaic3D基础模型,结合对比学习训练的3D编码器与轻量级掩码解码器,实现开放词汇3D语义与实例分割。在ScanNet200、Matterport3D和ScanNet++等任务上取得当前最佳性能,消融实验验证了大规模训练数据的有效性。

原文摘要 · Abstract (English)

We tackle open-vocabulary 3D scene understanding by introducing a novel data generation pipeline and training framework. Our method addresses three critical requirements for effective training: precise 3D region segmentation, comprehensive textual descriptions, and sufficient dataset scale. By leveraging state-of-the-art open-vocabulary image segmentation models and region-aware Vision-Language Models, we develop an automatic pipeline that generates high-quality 3D mask-text pairs. Applying this pipeline to multiple 3D scene datasets, we create Mosaic3D-5.6M, a dataset of over 30K annotated scenes with 5.6M mask-text pairs, significantly larger than existing datasets. Building upon this data, we propose Mosaic3D, a foundation model combining a 3D encoder trained with contrastive learning and a lightweight mask decoder for open-vocabulary 3D semantic and instance segmentation. Our approach achieves state-of-the-art results on open-vocabulary 3D semantic and instance segmentation tasks including ScanNet200, Matterport3D, and ScanNet++, with ablation studies validating the effectiveness of our large-scale training data.

3D分割开放词汇基础模型数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。