从单帧高清转播画面中精准定位运动员世界坐标。
A Top-Down Framework for Metric-Scale Athlete Localization from Single Broadcast Frames

- 自适应切片技术确保物体完整包含,减少边界切割误差。
- 在公开测试集上达到97.44的LocSim和0.9128的mAP,性能提升超21%。
- 适合高分辨率体育视频分析,对尺度变化鲁棒性强。
从单帧高清转播画面中实现运动员世界坐标精确定位极具挑战性,主要源于极端尺度差异。本文提出一种自顶向下的度量尺度定位框架。首先,提出边界感知自适应切片方法,基于粗略框预测迭代扩展切片边界,以语义引导方式确保目标完整包含,通过轻量级流程改进有效缓解边界切割伪影,且无需修改网络结构。该方法显著降低极端尺度变化下的召回率下降,使透视失真成为残差误差的主要来源。其次,将RTMPose-X架构改造为专用双关键点估计算法(骨盆与地面投影点),采用针对几何耦合点对优化的门控注意力单元,并通过相机标定后的射线投射实现2D投影点到世界坐标的确定性提升。在公开测试集上,本方法获得97.44的LocSim分数和0.9128的mAP,相比基线提升超过21%,建立了应对高分辨率尺度变异的鲁棒解决方案。
原文摘要 · Abstract (English)
Accurate world-coordinate localization of athletes from single-frame broadcast footage is inherently challenging due to extreme scale disparities in ultra-high-resolution imagery. In this paper, we propose a top-down framework for metric-scale athlete localization from a single calibrated frame. Our approach centers on three key contributions. First, we propose Boundary-Aware Adaptive Tiling, a semantics-guided extension of standard sliced inference. By iteratively expanding tile boundaries based on coarse bounding-box predictions, it systematically ensures full object containment, effectively mitigating boundary-splitting artifacts through a lightweight pipeline adaptation without architectural modifications. By substantially mitigating recall degradation under extreme scale variance, Boundary-Aware Adaptive Tiling enables us to isolate perspective distortion as the primary source of residual localization error. Second, we adapt the RTMPose-X architecture into a specialized two-keypoint estimator (pelvis and ground projection), employing a reformulated Gated Attention Unit optimized for this geometrically coupled point pair, and then deterministically lift the 2D ground projections into world coordinates via camera-calibrated ray casting. On the public test set, our method achieves a LocSim score of 97.44 and an mAP of 0.9128, outperforming the baseline by over 21 \% and establishing a robust solution for high-resolution scale variance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。