arXiv:2507.06662cs.CVcs.RO2025-07被引 7

融合视觉与文本信息,提升物体姿态估计的泛化能力。

MK-Pose: Category-Level Object Pose Estimation via Multimodal-Based Keypoint Learning

  • 结合图像、点云和文本描述,实现多模态关键点学习。
  • 在CAMERA25和REAL275上达到最优性能,无需形状先验。
  • 适合工业场景中跨类别、遮挡情况下的物体定位任务。

类别级物体姿态估计旨在不依赖具体实例先验的情况下,预测某类别内物体的姿态,对仓储自动化与制造场景至关重要。现有方法依赖于RGB图像或点云数据,常因物体遮挡及跨实例、跨类别泛化能力差而受限。本文提出一种基于多模态的关键点学习框架MK-Pose,融合RGB图像、点云与类别级文本描述。模型采用自监督关键点检测模块,结合注意力查询生成、软热图匹配与图关系建模;同时设计图增强特征融合模块,整合局部几何信息与全局上下文。在CAMERA25和REAL275数据集上评估,并在HouseCat6D数据集上测试跨数据集能力。结果表明,MK-Pose在无形状先验条件下,于IoU与平均精度指标上均优于现有最先进方法。代码将公开于https://github.com/yangyifanYYF/MK-Pose。

原文摘要 · Abstract (English)

Category-level object pose estimation, which predicts the pose of objects within a known category without prior knowledge of individual instances, is essential in applications like warehouse automation and manufacturing. Existing methods relying on RGB images or point cloud data often struggle with object occlusion and generalization across different instances and categories. This paper proposes a multimodal-based keypoint learning framework (MK-Pose) that integrates RGB images, point clouds, and category-level textual descriptions. The model uses a self-supervised keypoint detection module enhanced with attention-based query generation, soft heatmap matching and graph-based relational modeling. Additionally, a graph-enhanced feature fusion module is designed to integrate local geometric information and global context. MK-Pose is evaluated on CAMERA25 and REAL275 dataset, and is further tested for cross-dataset capability on HouseCat6D dataset. The results demonstrate that MK-Pose outperforms existing state-of-the-art methods in both IoU and average precision without shape priors. Codes will be released at \href{https://github.com/yangyifanYYF/MK-Pose}{https://github.com/yangyifanYYF/MK-Pose}.

姿态估计多模态关键点学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。