arXiv:2501.02831cs.CV2025-01AAAI

无需微调即可估计未知物体的6自由度姿态,利用2D与3D通用特征匹配。

Universal Features Guided Zero-Shot Category-Level Object Pose Estimation

  • 结合2D和3D通用特征建立语义对应关系,实现零样本姿态估计。
  • 在REAL275和Wild6D上对未见类别性能优于已有方法。
  • 适合机器人抓取、虚拟现实等需快速适应新物体的场景。

物体姿态估计在计算机视觉与机器人应用中至关重要,但面对未见类别的多样性仍具挑战。本文提出一种零样本方法,实现类别级6-DOF物体姿态估计,利用输入RGB-D图像的2D与3D通用特征建立基于语义相似性的对应关系,可扩展至未见类别而无需额外模型微调。方法首先通过高效2D通用特征在同类物体间寻找稀疏对应,获得初始粗略姿态;当姿态偏离较大时,2D特征对应退化,采用迭代策略优化姿态。随后,为解决同类物体间形状差异导致的姿态歧义,利用3D通用特征的密集对齐约束进一步精修粗略姿态。该方法在REAL275与Wild6D基准上对未见类别表现优于现有方法。

原文摘要 · Abstract (English)

Object pose estimation, crucial in computer vision and robotics applications, faces challenges with the diversity of unseen categories. We propose a zero-shot method to achieve category-level 6-DOF object pose estimation, which exploits both 2D and 3D universal features of input RGB-D image to establish semantic similarity-based correspondences and can be extended to unseen categories without additional model fine-tuning. Our method begins with combining efficient 2D universal features to find sparse correspondences between intra-category objects and gets initial coarse pose. To handle the correspondence degradation of 2D universal features if the pose deviates much from the target pose, we use an iterative strategy to optimize the pose. Subsequently, to resolve pose ambiguities due to shape differences between intra-category objects, the coarse pose is refined by optimizing with dense alignment constraint of 3D universal features. Our method outperforms previous methods on the REAL275 and Wild6D benchmarks for unseen categories.

姿态估计零样本学习通用特征6-DOF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。