arXiv:2506.19331cs.CV2025-06被引 10

用一句话描述3D场景中任意部分,实现精准语义分割

Segment Any 3D-Part in a Scene from a Sentence

  • 基于自然语言输入,实现3D场景中细粒度部件级分割
  • 在多个数据集上表现优异,具备强泛化能力
  • 提出首个大规模带密集部件标注的3D数据集3D-PU

本文旨在基于自然语言描述对3D场景中的任意部分进行分割,突破传统对象级3D理解的局限,解决数据与方法上的挑战。由于获取和标注成本高,现有数据集和方法大多局限于对象级理解。为克服数据与标注不足的问题,我们提出了3D-PU数据集——首个大规模带有密集部件标注的3D数据集,通过创新且低成本的合成3D场景构建方法生成细粒度部件级标注,为高级3D部件场景理解铺平道路。在方法层面,我们提出OpenPart3D,一种仅以3D输入为主的框架,有效应对部件级分割挑战。大量实验表明,该方法在开放词汇3D场景理解任务中表现出色,跨多个3D场景数据集具有强泛化能力。

原文摘要 · Abstract (English)

This paper aims to achieve the segmentation of any 3D part in a scene based on natural language descriptions, extending beyond traditional object-level 3D scene understanding and addressing both data and methodological challenges. Due to the expensive acquisition and annotation burden, existing datasets and methods are predominantly limited to object-level comprehension. To overcome the limitations of data and annotation availability, we introduce the 3D-PU dataset, the first large-scale 3D dataset with dense part annotations, created through an innovative and cost-effective method for constructing synthetic 3D scenes with fine-grained part-level annotations, paving the way for advanced 3D-part scene understanding. On the methodological side, we propose OpenPart3D, a 3D-input-only framework to effectively tackle the challenges of part-level segmentation. Extensive experiments demonstrate the superiority of our approach in open-vocabulary 3D scene understanding tasks at the part level, with strong generalization capabilities across various 3D scene datasets.

3D分割自然语言部件级

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。