arXiv:2412.01820cs.CV2024-12CVPR被引 32

构建首个超大规模足球视频数据集与专用视觉模型

Towards Universal Soccer Video Understanding

  • 打造1988场完整比赛的多模态数据集,自动标注
  • 新模型MatchVision在事件分类等任务上领先现有方法
  • 适合体育智能、视频理解研究者参考

作为全球广受欢迎的运动,足球吸引了全世界球迷的关注。本文致力于构建一个全面的多模态足球视频理解框架。具体贡献包括:(i) 提出SoccerReplay-1988,目前最大的多模态足球数据集,包含1988场完整比赛的视频与详细标注,采用自动化标注流程;(ii) 提出一种先进的足球专用视觉编码器MatchVision,利用足球视频中的时空信息,在多种下游任务中表现优异;(iii) 在事件分类、解说生成和多视角犯规识别上开展广泛实验与消融研究。MatchVision在所有任务中均达到当前最佳性能,显著优于现有模型,凸显所提数据与模型的优势。本工作有望为体育理解研究提供标准范式。

原文摘要 · Abstract (English)

As a globally celebrated sport, soccer has attracted widespread interest from fans all over the world. This paper aims to develop a comprehensive multi-modal framework for soccer video understanding. Specifically, we make the following contributions in this paper: (i) we introduce SoccerReplay-1988, the largest multi-modal soccer dataset to date, featuring videos and detailed annotations from 1,988 complete matches, with an automated annotation pipeline; (ii) we present an advanced soccer-specific visual encoder, MatchVision, which leverages spatiotemporal information across soccer videos and excels in various downstream tasks; (iii) we conduct extensive experiments and ablation studies on event classification, commentary generation, and multi-view foul recognition. MatchVision demonstrates state-of-the-art performance on all of them, substantially outperforming existing models, which highlights the superiority of our proposed data and model. We believe that this work will offer a standard paradigm for sports understanding research.

足球分析多模态视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。