从角落摄像头远距离精准还原被遮挡的手部3D姿态
REACH: Hand Pose Estimation from Room Corners

- 利用手体协同与多视角时序特征,基于Transformer建模
- 在50人多样动作数据集上实现高精度远距离手姿估计
- 适合无人干预的自然行为连续分析场景
我们提出一种新型3D手部姿态估计算法,可从房间角落固定摄像头在极低分辨率且频繁遮挡的视图中,远距离准确恢复人的手部形状与姿态。核心思路是充分利用手体协调性、时序动态变化及多视角观测信息。采用新型Transformer模型,通过视图级特征令牌间的相关性建模手与身体配置,并以自回归方式捕捉其时间一致性。为此构建首个大规模手姿数据集REACH(Room-Environment dataset Annotated with Chest cameras for Hand pose estimation),涵盖50名参与者在多种日常活动中的精确手部运动。为避免干扰自然动作,使用隐蔽胸贴摄像头标注。大量实验表明,所提方法REACH-Net在现有方法对比中表现优异,显著拓展了3D手部姿态估计在真实场景下持续人类行为分析的应用边界。
原文摘要 · Abstract (English)
We introduce a novel 3D hand pose estimator that can accurately recover the shape and pose of people's hands in a room from afar, typically from fixed cameras at room corners, in extremely low-resolution and frequently occluded views. Our key idea is to fully leverage hand-body coordination, its temporal progression, and multiview observations. We achieve this with a novel Transformer-based model, in which hand and body configurations are modeled through correlations between their visual features expressed as per-view tokens, and their temporal coordination is exploited in an autoregressive manner. We introduce a novel dataset, which we refer to as REACH, Room-Environment dataset Annotated with Chest cameras for Hand pose estimation, to train and test our method. REACH is a first-of-its-kind large-scale hand pose dataset that captures accurate hand movements of 50 participants across a wide variety of daily activities. In order to avoid interfering with natural movements while annotating the hands with accurate shape and pose, we leverage concealed chest cameras. Through extensive experiments, including comparative studies with existing methods, we show that our model, REACH-Net, achieves highly accurate 3D hand pose estimation from afar. These results broaden the horizon of 3D hand pose estimation, especially towards "in-the-wild" continuous human behavior analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。