用人体动作模型迁移学习,让2D手部姿态估计更准
TransHands: Repurposing Human Pose Encoders as Hand Pose Encoders

- 复用人体动作编码器,通过轻量适配模块对齐手部运动空间
- 在4种主流架构上均提升精度,尤其在自拍视角下表现优异
- 适合缺乏3D手部数据的实时应用,如手势交互与动作捕捉
从单目2D姿态恢复3D手部姿态仍具挑战,主要因高质量、多样化的3D标注手部数据集稀缺,而人体运动数据丰富。本文提出TransHands,一种不依赖主干网络的迁移学习框架,将大规模人体动作数据中学习到的运动先验迁移到手部领域。通过两阶段训练与微调策略,结合轻量级手部输入适配模块,使预训练的人体编码器能有效处理2D手部输入。我们在四种先进动作建模架构(包括基于Transformer、图结构和频域模型)上验证该方法。结果表明,人体运动先验可稳定跨架构迁移,在多种场景中一致提升精度,尤其在具有挑战性的自拍视角下表现突出,且适用于真实场景中的下游任务。
原文摘要 · Abstract (English)
Lifting 3D hand poses from 2D monocular representations remains challenging due to the limited availability of large-scale, diverse 3D-annotated hand datasets, in contrast to the abundance of human body motion data. We address this limitation by transferring motion representations learned from large body pose corpora to the hand domain. We introduce TransHands, a backbone-agnostic transfer learning framework that enables pre-trained human motion encoders to be effectively adapted for 3D hand pose estimation from 2D pose inputs. Rather than training hand-specific biomechanical models from scratch, TransHands combines a two-stage training and fine-tuning strategy with a lightweight hand-specific input adaptation module that aligns hand kinematics with the representation space learned for full-body motion. We evaluate TransHands across four state-of-the-art motion modeling architectures, including transformer-based, graph-based, and frequency- domain models. Results demonstrate that motion priors learned from body pose data transfer consistently across architectures, yielding consistent accuracy gains, strong cross-domain generalization, particularly in challenging egocentric settings, and applicability for downstream tasks in real-world contexts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。