基于视觉变压器的双手姿态估计算法,实现高精度动作捕捉与熟练度评估。
PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge
- 融合ViT与CNN的HP-ViT+架构,通过加权融合提升手部关键点定位精度。
- 手部姿态挑战赛取得8.31毫米平均误差,身体姿态挑战赛达11.25毫米误差。
- 首次将姿态估计技术扩展至熟练度评估,获顶尖表现(0.53准确率)。
本文介绍了我们团队(PCIE_EgoPose)在CVPR2025举办的EgoExo4D姿态与熟练度评估挑战赛中的解决方案。针对从RGB第一人称视频中估计21个3D手部关节这一复杂任务,该任务因细微动作和频繁遮挡而极具挑战性,我们提出了手部姿态视觉变换器(HP-ViT+)。该架构结合视觉变换器与卷积神经网络骨干网络,采用加权融合策略优化手部姿态预测。在EgoExo4D人体姿态挑战中,我们采用多模态时空特征融合策略,应对动态环境下的姿态估计难题。所提方法在手部姿态挑战中取得8.31毫米的PA-MPJPE,在身体姿态挑战中达到11.25毫米的MPJPE,均获得冠军。我们将姿态估计技术延伸至熟练度评估任务,利用基于变换器的核心架构,实现演示者熟练度评估的0.53顶级准确率,刷新当前最优结果。
原文摘要 · Abstract (English)
This report introduces our team's (PCIE_EgoPose) solutions for the EgoExo4D Pose and Proficiency Estimation Challenges at CVPR2025. Focused on the intricate task of estimating 21 3D hand joints from RGB egocentric videos, which are complicated by subtle movements and frequent occlusions, we developed the Hand Pose Vision Transformer (HP-ViT+). This architecture synergizes a Vision Transformer and a CNN backbone, using weighted fusion to refine the hand pose predictions. For the EgoExo4D Body Pose Challenge, we adopted a multimodal spatio-temporal feature integration strategy to address the complexities of body pose estimation across dynamic contexts. Our methods achieved remarkable performance: 8.31 PA-MPJPE in the Hand Pose Challenge and 11.25 MPJPE in the Body Pose Challenge, securing championship titles in both competitions. We extended our pose estimation solutions to the Proficiency Estimation task, applying core technologies such as transformer-based architectures. This extension enabled us to achieve a top-1 accuracy of 0.53, a SOTA result, in the Demonstrator Proficiency Estimation competition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。