用状态空间模型提升复杂遮挡下的3D手部姿态估计精度
HandMCM: Multi-modal Point Cloud-based Correspondence State Space Model for 3D Hand Pose Estimation

- 基于Mamba的状态空间模型,融合局部信息与关键点对应关系建模
- 在三个数据集上显著优于现有方法,尤其在严重遮挡场景下
- 适合需高精度手姿识别的AR/VR、人机交互等应用
3D手部姿态估计对增强现实等人机交互应用至关重要,但自遮挡及与物体交互导致的遮挡带来挑战。本文提出HandMCM,一种基于强大状态空间模型(Mamba)的新方法。通过引入局部信息注入/过滤模块和对应关系建模,该模型能有效学习不同遮挡场景下关键点的高度动态运动拓扑结构。同时,融合多模态图像特征以增强输入的鲁棒性与表征能力,实现更精准的手部姿态估计。在三个基准数据集上的实证评估表明,本模型显著优于当前最先进方法,尤其在严重遮挡场景下表现突出。结果证明了该方法在实际应用中提升3D手部姿态估计精度与可靠性的潜力。
原文摘要 · Abstract (English)
3D hand pose estimation that involves accurate estimation of 3D human hand keypoint locations is crucial for many human-computer interaction applications such as augmented reality. However, this task poses significant challenges due to self-occlusion of the hands and occlusions caused by interactions with objects. In this paper, we propose HandMCM to address these challenges. Our HandMCM is a novel method based on the powerful state space model (Mamba). By incorporating modules for local information injection/filtering and correspondence modeling, the proposed correspondence Mamba effectively learns the highly dynamic kinematic topology of keypoints across various occlusion scenarios. Moreover, by integrating multi-modal image features, we enhance the robustness and representational capacity of the input, leading to more accurate hand pose estimation. Empirical evaluations on three benchmark datasets demonstrate that our model significantly outperforms current state-of-the-art methods, particularly in challenging scenarios involving severe occlusions. These results highlight the potential of our approach to advance the accuracy and reliability of 3D hand pose estimation in practical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。