用状态空间模型提升单图双手交互三维重建精度与效率
VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image
- 引入状态空间模型融合局部全局特征,建模双手交互关系
- 在InterHand2.6M上实现MPJPE和MPVPE降低2-3%的突破
- 适合需要高精度手部交互重建的应用场景
理解双手交互对实现真实的3D姿态与形状重建至关重要。然而,现有方法在遮挡、外观模糊和计算效率方面仍面临挑战。为此,我们提出视觉马尔可夫双臂手交互网络(VM-BHINet),将状态空间模型(SSMs)引入手部重建,以增强交互建模能力并提升计算效率。核心组件——视觉马尔可夫交互特征提取块(VM-IFEBlock)结合了状态空间模型与局部、全局特征操作,实现了对手部交互的深度理解。在InterHand2.6M数据集上的实验表明,VM-BHINet将平均关节位置误差(MPJPE)和平均顶点位置误差(MPVPE)降低了2-3%,显著超越当前最优方法。
原文摘要 · Abstract (English)
Understanding bimanual hand interactions is essential for realistic 3D pose and shape reconstruction. However, existing methods struggle with occlusions, ambiguous appearances, and computational inefficiencies. To address these challenges, we propose Vision Mamba Bimanual Hand Interaction Network (VM-BHINet), introducing state space models (SSMs) into hand reconstruction to enhance interaction modeling while improving computational efficiency. The core component, Vision Mamba Interaction Feature Extraction Block (VM-IFEBlock), combines SSMs with local and global feature operations, enabling deep understanding of hand interactions. Experiments on the InterHand2.6M dataset show that VM-BHINet reduces Mean per-joint position error (MPJPE) and Mean per-vertex position error (MPVPE) by 2-3%, significantly surpassing state-of-the-art methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。