arXiv:2411.13550cs.CV2024-11ICCV被引 36

用2D模型自动生成海量3D部件数据,实现任意物体任意部件的零样本识别。

Find Any Part in 3D

  • 用2D基础模型构建数据引擎,自动生成1755倍于现有数据的部件类型
  • 零样本下mIoU提升260%,推理速度加快6至300倍
  • 适合需要开放世界3D部件分割的科研与工业应用

为何3D领域尚未出现基础模型?核心瓶颈在于数据稀缺。当前3D物体部件分割数据集规模小且多样性不足。我们证明可通过基于2D基础模型的数据引擎突破这一限制。该引擎可自动标注任意数量的物体部件,生成的部件类型总数达现有数据集总和的1755倍。利用此标注数据,仅通过简单的对比学习目标训练,即可获得一个开放世界的3D部件分割模型,能根据任意文本查询泛化到任意物体的任意部件。即使在零样本条件下,其性能仍优于在相应数据集上训练的传统方法,在mIoU上提升260%,同时推理速度提升6至300倍。缩放分析证实该泛化能力源于数据量的显著增长,凸显了数据引擎的关键作用。最后,为推动通用类别开放世界3D部件分割研究,我们发布了覆盖广泛物体与部件的新基准。项目网站:https://ziqi-ma.github.io/find3dsite/

原文摘要 · Abstract (English)

Why don't we have foundation models in 3D yet? A key limitation is data scarcity. For 3D object part segmentation, existing datasets are small in size and lack diversity. We show that it is possible to break this data barrier by building a data engine powered by 2D foundation models. Our data engine automatically annotates any number of object parts: 1755x more unique part types than existing datasets combined. By training on our annotated data with a simple contrastive objective, we obtain an open-world model that generalizes to any part in any object based on any text query. Even when evaluated zero-shot, we outperform existing methods on the datasets they train on. We achieve 260% improvement in mIoU and boost speed by 6x to 300x. Our scaling analysis confirms that this generalization stems from the data scale, which underscores the impact of our data engine. Finally, to advance general-category open-world 3D part segmentation, we release a benchmark covering a wide range of objects and parts. Project website: https://ziqi-ma.github.io/find3dsite/

3D分割开放世界数据生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。