用轻量模块融合手部细节与全身结构,提升3D人体姿态估计中手部精度。
Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator
- 通过条件调制模块融合预训练手部与全身模型特征。
- 手部关节与形状预测精度显著提升,整体姿态质量改善。
- 无需重训练即可实现高精度手部细节,适合需要精细手部建模的应用。
在3D全身姿态估计中,准确恢复手部姿态仍是一项重大挑战。根本原因在于监督缺口:全身姿态估计算法在包含有限手部多样性的全身体数据集上训练,而仅关注手部的估计算法虽在手指精细动作上表现优异,却缺乏全局身体上下文意识。为此,我们提出Hand4Whole++,一种模块化框架,利用预训练的全身与手部姿态估计算法的优势。引入CHAM(条件手部调制器),一个轻量级模块,通过预训练手部估计器提取的手部特征调制全身特征流。该调制使全身模型能预测既准确又符合上肢运动学结构的腕部朝向,且无需重新训练全身模型。同时,直接引入手部估计器预测的指关节角度和手部形状,并通过可微刚性对齐将其与全身网格对齐。该设计实现了全局一致的身体推理与细粒度的手部细节结合。大量实验表明,Hand4Whole++显著提升了手部精度,并改善了整体全身体姿态质量。
原文摘要 · Abstract (English)
Accurately recovering hand poses within the body context remains a major challenge in 3D whole-body pose estimation. This difficulty arises from a fundamental supervision gap: whole-body pose estimators are trained on full-body datasets with limited hand diversity, while hand-only estimators, trained on hand-centric datasets, excel at detailed finger articulation but lack global body awareness. To address this, we propose Hand4Whole++, a modular framework that leverages the strengths of both pre-trained whole-body and hand pose estimators. We introduce CHAM (Conditional Hands Modulator), a lightweight module that modulates the whole-body feature stream using hand-specific features extracted from a pre-trained hand pose estimator. This modulation enables the whole-body model to predict wrist orientations that are both accurate and coherent with the upper-body kinematic structure, without retraining the full-body model. In parallel, we directly incorporate finger articulations and hand shapes predicted by the hand pose estimator, aligning them to the full-body mesh via differentiable rigid alignment. This design allows Hand4Whole++ to combine globally consistent body reasoning with fine-grained hand detail. Extensive experiments demonstrate that Hand4Whole++ substantially improves hand accuracy and enhances overall full-body pose quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。