用拓扑先验和混合图融合提升复杂物体6D姿态估计精度
THE-Pose: Topological Prior with Hybrid Graph Fusion for Estimating Category-Level 6D Object Pose
- 引入图像域拓扑特征与点云融合,兼顾全局上下文与局部结构
- 在REAL275上相比3D-GC基线提升35.8%,超越现有最优7.2%
- 适合处理遮挡严重、类内差异大的物体姿态估计任务
类别级物体6D姿态估计需同时依赖全局上下文与局部结构以应对类内差异。现有基于3D图卷积的方法仅关注局部几何与深度信息,对复杂物体和视觉模糊敏感。为此,我们提出THE-Pose框架,通过表面嵌入引入拓扑先验,并设计混合图融合(HGF)模块,自适应融合图像域拓扑特征与点云特征,无缝连接2D上下文与3D结构。该融合特征在未见或复杂物体上仍具稳定性,即使在显著遮挡下也表现稳健。在REAL275数据集上的大量实验表明,THE-Pose相较3D-GC基线(HS-Pose)提升35.8%,且优于此前最先进方法7.2%。代码已开源:https://github.com/EHxxx/THE-Pose
原文摘要 · Abstract (English)
Category-level object pose estimation requires both global context and local structure to ensure robustness against intra-class variations. However, 3D graph convolution (3D-GC) methods only focus on local geometry and depth information, making them vulnerable to complex objects and visual ambiguities. To address this, we present THE-Pose, a novel category-level 6D pose estimation framework that leverages a topological prior via surface embedding and hybrid graph fusion. Specifically, we extract consistent and invariant topological features from the image domain, effectively overcoming the limitations inherent in existing 3D-GC based methods. Our Hybrid Graph Fusion (HGF) module adaptively integrates the topological features with point-cloud features, seamlessly bridging 2D image context and 3D geometric structure. These fused features ensure stability for unseen or complicated objects, even under significant occlusions. Extensive experiments on the REAL275 dataset show that THE-Pose achieves a 35.8% improvement over the 3D-GC baseline (HS-Pose) and surpasses the previous state-of-the-art by 7.2% across all key metrics. The code is avaialbe on https://github.com/EHxxx/THE-Pose
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。