用2D数据预训练,提升单目人体动作恢复的精度与泛化能力
Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
- 将3D动作恢复视为多视角合成,分两阶段训练:先2D预训练,再3D微调
- 在野外数据集上实现更真实的相机空间动作和精准的物理世界定位
- 新姿态表示法分离局部动作与全局运动,加速收敛,适合真实场景应用
真实交互中的人体动作恢复需精确动作细节与度量尺度轨迹。从单目输入恢复绝对人体姿态虽具可行性,但仍面临两大挑战:(1) 模型依赖受限环境的3D训练数据,泛化能力差;(2) 单目观测难以准确估计度量尺度。本文提出Mocap-2-to-3框架,通过利用大规模2D数据增强3D动作恢复,区别于以往方法,直接从单目输入恢复绝对姿态。为有效利用2D数据中的动作先验与多样性,我们将3D动作恢复重构为多视角合成过程,并采用两阶段训练:先在大量2D数据上预训练单视角扩散模型,再在3D数据上进行多视角微调,融合强先验与几何约束。此外,为恢复绝对姿态,引入一种新型人体动作表示,解耦局部姿态与全局运动学习,同时编码地面几何先验以加速收敛,显著提升物理空间定位精度。在野外基准测试中,本方法在相机空间动作真实度与世界坐标定位方面均优于现有最优方法,且具备强大泛化能力。
原文摘要 · Abstract (English)
Human motion recovery for real-world interaction demands both precise action details and metric-scale trajectories. Recovering absolute human pose from monocular input presents a viable solution, but faces two main challenges: (1) models' reliance on 3D training data from constrained environments limits their out-of-distribution generalization; and (2) the inherent difficulty of estimating metric-scale poses from monocular observations. This paper introduces Mocap-2-to-3, a novel framework that differs from prior HMR methods by recovering absolute poses from monocular input and leveraging abundant 2D data to enhance 3D motion recovery. To effectively utilize the action priors and diversity in large-scale 2D datasets, we reformulate 3D motion as a multi-view synthesis process and divide the training into two stages: a single-view diffusion model is first pre-trained on extensive 2D data, followed by multi-view fine-tuning on 3D data, thus achieving a combination of strong priors and geometric constraints. Furthermore, to recover absolute poses, we introduce a novel human motion representation that decouples the learning of local pose and global movements, while encoding ground geometric priors to accelerate convergence, thereby yielding more precise positioning in the physical world. Experiments on in-the-wild benchmarks show that our method outperforms state-of-the-art approaches in both camera-space motion realism and world-grounded human positioning, while exhibiting strong generalization capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。