首个足球视觉基础模型,统一处理从识别到推理的多种任务。
SoccerMaster: A Vision Foundation Model for Soccer Understanding
- 用多任务监督预训练构建统一框架,整合多样足球理解任务。
- 自动生成空间标注,融合多个数据集形成大规模预训练资源。
- 在下游任务中超越专用模型,适合需要多任务理解的场景。
足球理解因领域特异性复杂性和独特挑战近年受到越来越多关注。与以往依赖孤立、任务特定专家模型的研究不同,本文提出SoccerMaster——首个针对足球领域的视觉基础模型,通过监督多任务预训练,在单一框架内统一处理从细粒度感知(如运动员检测与识别)到高层语义推理(如事件分类)的多样化任务。具体贡献包括:(i) 提出SoccerMaster,首个基于监督多任务预训练的足球专用视觉基础模型;(ii) 开发自动化数据整理流水线SoccerFactory,生成可扩展的空间标注,并整合多个现有足球视频数据集,构建全面的多任务预训练数据资源;(iii) 通过广泛评估证明,SoccerMaster在多种下游任务中持续优于任务特定专家模型,展现出卓越的泛化能力与优势。相关数据、代码与模型将公开发布。
原文摘要 · Abstract (English)
Soccer understanding has recently garnered growing research interest due to its domain-specific complexity and unique challenges. Unlike prior works that typically rely on isolated, task-specific expert models, this work aims to propose a unified model to handle diverse soccer visual understanding tasks, ranging from fine-grained perception (e.g., athlete detection and identification) to high-level semantic reasoning (e.g., event classification). Concretely, our contributions are threefold: (i) we present SoccerMaster, the first soccer-specific vision foundation model that unifies diverse tasks within a single framework via supervised multi-task pretraining; (ii) we develop an automated data curation pipeline, SoccerFactory, to generate scalable spatial annotations, and integrate multiple existing soccer video datasets as a comprehensive pretraining data resource for multi-task pretraining; and (iii) we conduct extensive evaluations demonstrating that SoccerMaster consistently outperforms task-specific expert models across diverse downstream tasks, highlighting its breadth and superiority. The data, code, and model will be publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。