arXiv:2505.09455cs.CV2025-05被引 6

用足球战术语言提升视频动作检测的准确率

Beyond Pixels: Leveraging the Language of Soccer to Improve Spatio-Temporal Action Detection in Broadcast Videos

  • 引入足球比赛语境与团队动态建模,优化动作预测
  • 在低置信度场景下同时提升精确率与召回率
  • 适合需要高覆盖性足球分析的场景

当前最先进的时空动作检测(STAD)方法在从广播视频中提取足球事件方面表现良好。但在需全面覆盖赛事事件的高召回、低精度模式下,其缺乏上下文理解能力导致大量误报。本文通过在比赛层面进行推理,引入去噪序列转换任务,将噪声大、无上下文的以球员为中心的预测与清晰的比赛状态信息结合,使用基于Transformer的编码器-解码器模型联合建模长时间依赖关系和团队级动态。该方法利用足球的战术规律与球员间依赖关系,生成更准确的动作序列。实验表明,该方法在低置信度情况下显著提升精度与召回率,实现更可靠的事件提取,可有效补充现有基于像素的方法。

原文摘要 · Abstract (English)

State-of-the-art spatio-temporal action detection (STAD) methods show promising results for extracting soccer events from broadcast videos. However, when operated in the high-recall, low-precision regime required for exhaustive event coverage in soccer analytics, their lack of contextual understanding becomes apparent: many false positives could be resolved by considering a broader sequence of actions and game-state information. In this work, we address this limitation by reasoning at the game level and improving STAD through the addition of a denoising sequence transduction task. Sequences of noisy, context-free player-centric predictions are processed alongside clean game state information using a Transformer-based encoder-decoder model. By modeling extended temporal context and reasoning jointly over team-level dynamics, our method leverages the "language of soccer" - its tactical regularities and inter-player dependencies - to generate "denoised" sequences of actions. This approach improves both precision and recall in low-confidence regimes, enabling more reliable event extraction from broadcast video and complementing existing pixel-based methods.

动作检测足球分析Transformer上下文建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。