arXiv:2501.06785cs.CVcs.CL2025-01被引 5

构建200类3D物体细粒度部件与材质数据集,提升模型组合理解能力

3DCoMPaT200: Language-Grounded Compositional Understanding of Parts and Materials of 3D Shapes

  • 构建200类3D物体数据集,部件种类达1031个,材质293类
  • 提出文本引导的部件形状检索任务,多部件描述时模型性能显著提升
  • 适合研究3D视觉理解、生成与机器人交互的开发者使用

人类与机器人在环境中的导航和交互依赖于对3D物体的部件级理解。现有部件级3D理解数据集类别有限,如ShapeNet-Part和PartNet分别仅包含16和24个类别,3DCoMPaT仅有42个类别。为推动更丰富、细粒度的部件级3D理解,本文提出3DCoMPaT200,一个大规模数据集,涵盖200个物体类别,较3DCoMPaT增加约5倍物体词汇量,部件类别增加约4倍。3DCoMPaT200包含1,031个细粒度部件类别和293种不同材质类别,支持部件与材质的组合应用。为应对组合式3D建模的复杂性,我们提出一种新任务:基于文本描述的组合部件形状检索,采用ULIP作为强3D基础模型。该任务评估模型在给定1、3或6个部件文本描述下的形状检索性能。结果表明,随着风格组合数量增加,模型性能提升,凸显该数据集在增强模型组合理解能力方面的有效性。代码与数据见http://github.com/3DCoMPaT200/3DCoMPaT200。

原文摘要 · Abstract (English)

Understanding objects in 3D at the part level is essential for humans and robots to navigate and interact with the environment. Current datasets for part-level 3D object understanding encompass a limited range of categories. For instance, the ShapeNet-Part and PartNet datasets only include 16, and 24 object categories respectively. The 3DCoMPaT dataset, specifically designed for compositional understanding of parts and materials, contains only 42 object categories. To foster richer and fine-grained part-level 3D understanding, we introduce 3DCoMPaT200, a large-scale dataset tailored for compositional understanding of object parts and materials, with 200 object categories with $\approx$5 times larger object vocabulary compared to 3DCoMPaT and $\approx$ 4 times larger part categories. Concretely, 3DCoMPaT200 significantly expands upon 3DCoMPaT, featuring 1,031 fine-grained part categories and 293 distinct material classes for compositional application to 3D object parts. Additionally, to address the complexities of compositional 3D modeling, we propose a novel task of Compositional Part Shape Retrieval using ULIP to provide a strong 3D foundational model for 3D Compositional Understanding. This method evaluates the model shape retrieval performance given one, three, or six parts described in text format. These results show that the model's performance improves with an increasing number of style compositions, highlighting the critical role of the compositional dataset. Such results underscore the dataset's effectiveness in enhancing models' capability to understand complex 3D shapes from a compositional perspective. Code and Data can be found at http://github.com/3DCoMPaT200/3DCoMPaT200

3D理解部件识别组合推理数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。