用改进的SMPLest-X和光流追踪,提升足球比赛3D人体姿态估计精度。
SMART: SMPLest-X Mesh Adaptation and RAFT Tracking for Soccer Pose Estimation
- 分层片段微调+多任务深度监督+广播数据增强,优化人体模型
- 验证集得分0.647,比基线降低38.6%;测试集全局平均误差0.324米
- 适合体育视频分析、动作捕捉与实时追踪场景
我们提出的方法参与了2026年FIFA骨骼追踪挑战,旨在从转播视频中估计足球运动员在三维世界空间中的姿态。方法基于对SMPLest-X(ViT-H,687M参数)进行分层片段微调、多任务深度监督以及广播数据增强,并结合RAFT密集光流相机追踪器、脚底平面锚定和两阶段时间平滑。在验证集上,SMART得分0.647,相比FIFA基线1.053降低38.6%;在预留测试集上得分为0.593(全局平均关节位置误差:0.324米,局部平均误差:0.054米)。
原文摘要 · Abstract (English)
We present our approach to the FIFA Skeletal Tracking Challenge 2026, which requires estimating 3D world-space poses of soccer players from broadcast video. Our method finetunes SMPLest-X (ViT-H, 687 M parameters) via a stratified clip split, multi-task depth supervision, and broadcast augmentation, paired with a RAFT dense optical flow camera tracker, foot-plane anchoring, and two-pass temporal smoothing. Against the FIFA baseline score of 1.053 on the validation set, SMART achieves 0.647, a 38.6% improvement; on the held-out test set, SMART scores 0.593 (Global MPJPE: 0.324 m, Local MPJPE: 0.054 m).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。