通过重建全局语义形状提升物体姿态估计,尤其在部分可见时表现更优。
GCE-Pose: Global Context Enhancement for Category-level Object Pose Estimation
- 先重建类别整体形状再融合局部观测特征
- 在HouseCat6D和NOCS-REAL275上显著超越现有方法
- 适合处理遮挡严重的真实场景物体姿态估计
模型无关的类别级物体姿态估计面临跨实例泛化上下文特征提取难题。现有方法依赖基础特征捕捉语义与几何信息,但在部分可见时失效。本文提出GCE-Pose,采用‘先完整后聚合’策略,利用类别先验增强姿态估计。通过提出的语义形状重建(SSR)模块,将未见的局部RGB-D对象通过学习的深度线性形状模型变形为类别特有3D语义原型,实现全局几何与语义重建。进一步设计全局上下文增强(GCE)特征融合模块,有效融合部分观测与重建的全局上下文特征。大量实验证明,该全局上下文先验及GCE模块显著提升性能,在挑战性真实数据集HouseCat6D与NOCS-REAL275上优于现有方法。
原文摘要 · Abstract (English)
A key challenge in model-free category-level pose estimation is the extraction of contextual object features that generalize across varying instances within a specific category. Recent approaches leverage foundational features to capture semantic and geometry cues from data. However, these approaches fail under partial visibility. We overcome this with a first-complete-then-aggregate strategy for feature extraction utilizing class priors. In this paper, we present GCE-Pose, a method that enhances pose estimation for novel instances by integrating category-level global context prior. GCE-Pose performs semantic shape reconstruction with a proposed Semantic Shape Reconstruction (SSR) module. Given an unseen partial RGB-D object instance, our SSR module reconstructs the instance's global geometry and semantics by deforming category-specific 3D semantic prototypes through a learned deep Linear Shape Model. We further introduce a Global Context Enhanced (GCE) feature fusion module that effectively fuses features from partial RGB-D observations and the reconstructed global context. Extensive experiments validate the impact of our global context prior and the effectiveness of the GCE fusion module, demonstrating that GCE-Pose significantly outperforms existing methods on challenging real-world datasets HouseCat6D and NOCS-REAL275. Our project page is available at https://colin-de.github.io/GCE-Pose/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。