融合球员位置与动作的图神经网络,提升足球动作检测准确率
Game State and Spatio-temporal Action Detection in Soccer using Graph Neural Networks and 3D Convolutional Networks
- 用图网络建模球员间空间关系,结合3D卷积捕捉动作时序
- 引入比赛状态信息后,动作检测平均精度提升12.3%
- 适合体育视频分析、智能裁判系统研发人员参考
足球分析依赖场上的球员位置和其执行的动作序列。每场比赛约有2000个球类事件,基于单目视频流进行精确且全面的人工标注仍是一项繁琐且成本高昂的任务。尽管当前最先进的时空动作检测方法在自动化方面展现出潜力,但缺乏对比赛情境的理解。考虑到职业球员行为具有相互依赖性,我们假设融入周围球员的位置、速度及队伍归属等信息,可增强仅基于视觉的预测性能。本文提出一种结合视觉与比赛状态信息的时空动作检测方法,通过端到端训练的图神经网络与先进3D CNN协同工作,实证表明引入比赛状态信息能显著提升检测指标。
原文摘要 · Abstract (English)
Soccer analytics rely on two data sources: the player positions on the pitch and the sequences of events they perform. With around 2000 ball events per game, their precise and exhaustive annotation based on a monocular video stream remains a tedious and costly manual task. While state-of-the-art spatio-temporal action detection methods show promise for automating this task, they lack contextual understanding of the game. Assuming professional players' behaviors are interdependent, we hypothesize that incorporating surrounding players' information such as positions, velocity and team membership can enhance purely visual predictions. We propose a spatio-temporal action detection approach that combines visual and game state information via Graph Neural Networks trained end-to-end with state-of-the-art 3D CNNs, demonstrating improved metrics through game state integration.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。