基于世界杯影像构建超大规模3D人体姿态数据集,支持多视角精准追踪。
WorldPose: A World Cup Dataset for Global 3D Human Pose Estimation
- 利用多固定与移动摄像机,实现1.75英亩区域内的高精度3D姿态重建。
- 包含超过80段视频、250万帧3D姿态,总运动距离超120公里。
- 适合体育分析、多人全局姿态估计等研究,开源标注与摄像机参数。
我们提出WorldPose,一个面向真实场景下多人全局姿态估计的新数据集,内容来自2022年FIFA世界杯比赛画面。与以往聚焦单人或室内环境的数据集不同,本次赛事部署了多个固定与移动摄像机,覆盖多个体育场。利用高清多视角静态设置,我们实现了超过1.75英亩区域内的球员3D姿态与运动的高精度恢复。进一步,结合球员动作与场内标记点,校准了移动转播摄像机参数。最终数据集包含80余段序列,约250万帧3D姿态,总运动距离超过120公里。我们对当前最先进的全局姿态估计方法进行了深入分析,实验表明WorldPose显著挑战现有技术,推动该领域及体育分析等方向的研究。所有姿态标注(SMPL格式)、转播摄像机参数与视频素材将用于学术研究开放共享。
原文摘要 · Abstract (English)
We present WorldPose, a novel dataset for advancing research in multi-person global pose estimation in the wild, featuring footage from the 2022 FIFA World Cup. While previous datasets have primarily focused on local poses, often limited to a single person or in constrained, indoor settings, the infrastructure deployed for this sporting event allows access to multiple fixed and moving cameras in different stadiums. We exploit the static multi-view setup of HD cameras to recover the 3D player poses and motions with unprecedented accuracy given capture areas of more than 1.75 acres. We then leverage the captured players' motions and field markings to calibrate a moving broadcasting camera. The resulting dataset comprises more than 80 sequences with approx 2.5 million 3D poses and a total traveling distance of over 120 km. Subsequently, we conduct an in-depth analysis of the SOTA methods for global pose estimation. Our experiments demonstrate that WorldPose challenges existing multi-person techniques, supporting the potential for new research in this area and others, such as sports analysis. All pose annotations (in SMPL format), broadcasting camera parameters and footage will be released for academic research purposes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。