让3D物体分割理解功能角色,像人一样用标准视角分析形状
CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation through LLM-Guided Canonical Spatial Modeling
- 用大模型引导构建跨类别的标准空间数据集,发现物体的通用空间规律
- 通过双分支结构将不同姿态的物体统一映射到标准空间,稳定提取部件语义
- 适合做开放世界3D理解、通用物体识别与具身智能相关研究
开放世界提示式3D语义分割仍因语义依赖输入传感器坐标而脆弱。人类则通过标准空间中的功能角色理解物体——翅膀横向延伸,把手向外突出,腿从下方支撑。心理物理学研究表明,我们会将物体在脑中旋转至标准朝向以揭示其功能。为填补这一差距,我们提出CoSMo3D,通过数据驱动学习隐式标准参考系,实现对标准空间的感知。我们利用大语言模型引导,在200个类别间进行类别内与跨类别对齐,构建统一的标准空间数据集,揭示跨类别的标准空间规律。通过双分支架构实现模型内标准性:标准图锚定与标准框校准,将姿态变化和对称性压缩至稳定的标准嵌入表示。该从输入姿态空间到标准嵌入空间的转变,显著提升了部件语义的稳定性与可迁移性。实验表明,CoSMo3D在开放世界提示式3D分割任务上达到新基准。
原文摘要 · Abstract (English)
Open-world promptable 3D semantic segmentation remains brittle as semantics are inferred in the input sensor coordinates. Yet, humans, in contrast, interpret parts via functional roles in a canonical space -- wings extend laterally, handles protrude to the side, and legs support from below. Psychophysical evidence shows that we mentally rotate objects into canonical frames to reveal these roles. To fill this gap, we propose \methodName{}, which attains canonical space perception by inducing a latent canonical reference frame learned directly from data. By construction, we create a unified canonical dataset through LLM-guided intra- and cross-category alignment, exposing canonical spatial regularities across 200 categories. By induction, we realize canonicality inside the model through a dual-branch architecture with canonical map anchoring and canonical box calibration, collapsing pose variation and symmetry into a stable canonical embedding. This shift from input pose space to canonical embedding yields far more stable and transferable part semantics. Experimental results show that \methodName{} establishes new state of the art in open-world promptable 3D segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。