无需标注对应关系,通过形状先验实现物体级3D位置对齐。
Category-Level 3D Correspondence in Camera Space via Morphable Object Priors

- 基于可变形物体先验,从单图预测跨实例一致的3D位置。
- 在50类17.8万张图像上达到新最好性能,关键点对齐精度提升显著。
- 适合做机器人感知、虚拟现实中的物体语义理解研究者。
从图像理解3D物体是机器人与AR/VR应用的基础。尽管当前方法在类别级位姿估计上取得进展,但现有表征仍无法捕捉细粒度语义,难以支持对物体部件、功能及交互的推理。本文研究相机空间中的类别级3D对应问题——仅凭单张图像预测跨实例保持一致的3D位置,并证明该能力可通过学习共享的可变形物体先验,在无显式对应监督下自然涌现。为此,我们构建了首个大规模基准HouseCorr3D,包含50类家用物品、280个独立实例、17.8万张图像及直接标注于CAD模型上的3D关键点;其关键创新在于提供遮挡区域的非模态对应标签与显式对称性标注,弥补现有数据集缺陷。进一步提出Morpheus方法,通过解耦标准形状、形变与物体位姿,学习类别级可变形形状先验。借助这一共享的标准基底,语义有意义的相机空间3D对应关系得以隐式生成。在HouseCorr3D上,该方法达到新最优表现,证明语义3D理解可在无直接对应监督下实现。数据与代码已开源。
原文摘要 · Abstract (English)
Understanding 3D objects from images is fundamental to robotics and AR/VR applications. While recent work has made progress in category-level pose estimation, current representations fail to capture the fine-grained semantics needed for reasoning about object parts, functions, and interactions. In this work, we study category-level 3D correspondence in camera space -- predicting, from a single image, 3D locations that remain consistent across instances within a category -- and show that it can emerge without explicit correspondence supervision by learning a shared morphable object prior. To enable research in this direction, we introduce HouseCorr3D, the first large-scale benchmark for monocular category-level 3D correspondence with 178k images across 50 household object categories, 280 unique instances, and 3D keypoint annotations directly on CAD models. Crucially, HouseCorr3D provides amodal correspondence labels for occluded regions and explicit symmetry annotations, addressing key limitations of existing datasets. We further propose Morpheus, a method that learns morphable category-level shape priors by disentangling canonical shape, deformation, and object pose. Through this shared canonical grounding, semantically meaningful 3D correspondences in camera space emerge implicitly. These emerging 3D correspondences set a new state of the art on HouseCorr3D, demonstrating that semantic 3D object understanding can arise without direct correspondence supervision. Data and code are publicly available at https://github.com/GenIntel/HouseCorr3D.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。