融合鲁棒姿态估计与扩散轨迹先验,提升视频中人体动作重建精度。
RopeTP: Global Human Motion Recovery via Integrating Robust Pose Estimation with Diffusion Trajectory Prior
- 分层注意力机制增强上下文感知,改善遮挡部位姿态推断。
- 在两个基准数据集上超越现有方法,尤其在遮挡场景表现优异。
- 无需复杂优化,直接生成自然稳定的3D人体运动轨迹,适合实际应用。
我们提出RopeTP,一种结合鲁棒姿态估计与扩散轨迹先验的新型框架,用于从视频中重建全局人体运动。其核心是分层注意力机制,显著提升上下文感知能力,有助于准确推断被遮挡的身体部位。该机制通过利用可见解剖结构间的关系,增强局部姿态估计的鲁棒性,从而实现精确稳定的全局轨迹重建。此外,RopeTP引入扩散轨迹模型,从局部姿态序列预测出符合真实人类运动模式的动态轨迹,确保生成结果在时间上自然连贯,提升3D人体动作重建的真实感与稳定性。大量实验验证表明,RopeTP在两个基准数据集上均优于当前主流方法,尤其在遮挡场景下表现突出。同时,它无需依赖SLAM进行初始相机估计或复杂优化,仍能输出更准确、更真实的运动轨迹。
原文摘要 · Abstract (English)
We present RopeTP, a novel framework that combines Robust pose estimation with a diffusion Trajectory Prior to reconstruct global human motion from videos. At the heart of RopeTP is a hierarchical attention mechanism that significantly improves context awareness, which is essential for accurately inferring the posture of occluded body parts. This is achieved by exploiting the relationships with visible anatomical structures, enhancing the accuracy of local pose estimations. The improved robustness of these local estimations allows for the reconstruction of precise and stable global trajectories. Additionally, RopeTP incorporates a diffusion trajectory model that predicts realistic human motion from local pose sequences. This model ensures that the generated trajectories are not only consistent with observed local actions but also unfold naturally over time, thereby improving the realism and stability of 3D human motion reconstruction. Extensive experimental validation shows that RopeTP surpasses current methods on two benchmark datasets, particularly excelling in scenarios with occlusions. It also outperforms methods that rely on SLAM for initial camera estimates and extensive optimization, delivering more accurate and realistic trajectories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。