arXiv:2604.16954cs.CV2026-04

通过拓扑感知与语义建模提升物体姿态估计泛化能力

TSM-Pose: Topology-Aware Learning with Semantic Mamba for Category-Level Object Pose Estimation

  • 引入拓扑提取器捕捉点云全局结构,融合局部几何特征
  • 基于Mamba的语义聚合器增强关键点表达,建模长距离依赖
  • 在三个数据集上超越现有方法,适合机器人抓取等场景

类别级物体姿态估计是具身智能的基础,但对未见实例的鲁棒泛化仍具挑战。现有方法多依赖简单特征提取与聚合,难以捕捉类别共享的拓扑结构并进行语义关键点建模,限制了泛化性能。为此,我们提出TSM-Pose框架:引入拓扑提取器,捕捉点云全局拓扑表征,并融入局部几何特征,实现稳健的类别级结构表示;同时设计基于Mamba的全局语义聚合器,将语义先验注入关键点以增强表达力,并采用多个TwinMamba模块建模长程依赖,实现更有效的全局特征聚合。在REAL275、CAMERA25和HouseCat6D三个基准数据集上的大量实验表明,TSM-Pose优于现有最先进方法。

原文摘要 · Abstract (English)

Category-level object pose estimation is fundamental for embodied intelligence, yet achieving robust generalization to unseen instances remains challenging. However, existing methods mainly rely on simple feature extraction and aggregation, which struggle to capture category-shared topological structures and conduct semantic keypoint modeling, limiting their generalization. To address these, we propose a \textbf{T}opology-Aware Learning with \textbf{S}emantic \textbf{M}amba for Category-Level \textbf{P}ose Estimation framework (TSM-Pose). Specifically, we introduce a Topology Extractor to capture the global topological representation of the point cloud, which is integrated into local geometry features and enables robust category-level structural representation. Simultaneously, we propose a Mamba-based Global Semantic Aggregator that injects semantics priors into keypoints to enhance their expressiveness and leverages multiple TwinMamba blocks to model long-range dependencies for more effective global feature aggregation. Extensive experiments on three benchmark datasets (REAL275, CAMERA25, and HouseCat6D) demonstrate that TSM-Pose outperforms existing state-of-the-art methods.

姿态估计点云处理Mamba拓扑建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。