用单路直播视频实现足球场多目标检测追踪,让业余球队也能用专业级数据分析
A Computer Vision Framework for Multi-Class Detection and Tracking in Soccer Broadcast Footage
- 基于YOLO与ByteTrack构建端到端视觉系统,识别球员、裁判、门将和球
- 球员追踪精度高,召回率与mAP50表现良好,但球体检测仍是难点
- 无需专业设备即可获取球员位置数据,适合高校、青训和业余俱乐部
拥有昂贵多相机系统或GPS追踪设备的俱乐部能获得竞争优势,而低预算球队往往难以获取类似数据。本文探讨能否通过单一摄像机的计算机视觉流程,直接从标准转播画面中提取此类信息。本研究开发了一个端到端系统,结合YOLO目标检测器与ByteTrack跟踪算法,实现对比赛中球员、裁判、守门员及球的全程识别与追踪。实验结果表明,该流程在球员和官员的检测与追踪上表现优异,具备高精度、高召回率和良好的mAP50指标,但球体检测仍为主要挑战。尽管存在局限,研究证明人工智能可从单一路直播摄像头中提取有意义的球员级空间信息。该方法降低了对专用硬件的依赖,使高校、青训机构及业余俱乐部得以采用此前仅专业团队可用的可扩展数据驱动分析手段,凸显了低成本计算机视觉在足球分析中的潜力。
原文摘要 · Abstract (English)
Clubs with access to expensive multi-camera setups or GPS tracking systems gain a competitive advantage through detailed data, whereas lower-budget teams are often unable to collect similar information. This paper examines whether such data can instead be extracted directly from standard broadcast footage using a single-camera computer vision pipeline. This project develops an end-to-end system that combines a YOLO object detector with the ByteTrack tracking algorithm to identify and track players, referees, goalkeepers, and the ball throughout a match. Experimental results show that the pipeline achieves high performance in detecting and tracking players and officials, with strong precision, recall, and mAP50 scores, while ball detection remains the primary challenge. Despite this limitation, our findings demonstrate that AI can extract meaningful player-level spatial information from a single broadcast camera. By reducing reliance on specialized hardware, the proposed approach enables colleges, academies, and amateur clubs to adopt scalable, data-driven analysis methods previously accessible only to professional teams, highlighting the potential for affordable computer vision-based soccer analytics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。