arXiv:2512.02727cs.CVcs.AI2025-12中稿 · WACV 2026被引 2

用可变形状态空间模型提升遮挡下3D手部姿态估计精度

DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions

  • 引入可变形状态扫描机制,增强局部特征与全局上下文关联
  • 在5个数据集上超越现有模型,精度显著提升且推理速度媲美ResNet-50
  • 特别适合复杂交互场景中手部遮挡严重的3D姿态估计任务

日常手部交互常面临严重遮挡问题,如双手重叠,这对3D手部姿态估计(HPE)的鲁棒特征学习提出挑战。当前方法多依赖ResNet提取特征,其卷积神经网络的归纳偏置难以有效建模全局上下文。为此,本文提出基于状态空间模型(Mamba)的可变形状态空间框架——DF-Mamba,用于3D HPE中的视觉特征提取。通过Mamba的选通状态建模与提出的可变形状态扫描,该方法在卷积后特征上聚合图像信息,同时选择性保留代表全局上下文的有效线索。实验在五个差异较大的数据集上进行,涵盖单手/双手、手-物体交互及RGB/深度输入场景。DF-Mamba在所有数据集上均优于最新图像主干网络(如VMamba、Spatial-Mamba),达到当前最优性能,且推理速度与ResNet-50相当。

原文摘要 · Abstract (English)

Modeling daily hand interactions often struggles with severe occlusions, such as when two hands overlap, which highlights the need for robust feature learning in 3D hand pose estimation (HPE). To handle such occluded hand images, it is vital to effectively learn the relationship between local image features (e.g., for occluded joints) and global context (e.g., cues from inter-joints, inter-hands, or the scene). However, most current 3D HPE methods still rely on ResNet for feature extraction, and such CNN's inductive bias may not be optimal for 3D HPE due to its limited capability to model the global context. To address this limitation, we propose an effective and efficient framework for visual feature extraction in 3D HPE using recent state space modeling (i.e., Mamba), dubbed Deformable Mamba (DF-Mamba). DF-Mamba is designed to capture global context cues beyond standard convolution through Mamba's selective state modeling and the proposed deformable state scanning. Specifically, for local features after convolution, our deformable scanning aggregates these features within an image while selectively preserving useful cues that represent the global context. This approach significantly improves the accuracy of structured 3D HPE, with comparable inference speed to ResNet-50. Our experiments involve extensive evaluations on five divergent datasets including single-hand and two-hand scenarios, hand-only and hand-object interactions, as well as RGB and depth-based estimation. DF-Mamba outperforms the latest image backbones, including VMamba and Spatial-Mamba, on all datasets and achieves state-of-the-art performance.

3D姿态估计状态空间模型手部交互可变形扫描

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。