arXiv:2512.19528cs.CV2025-12被引 1

用多模态融合分析足球比赛,无需依赖精确球迹追踪。

Multi-Modal Soccer Scene Analysis with Masked Pre-Training

  • 融合球员轨迹、身份和图像块,通过时序变换器建模动态
  • 在真实顶级联赛数据上显著提升球轨迹、状态与持球者识别准确率
  • 提出专属遮蔽策略,避免视觉特征过拟合,适合体育视频分析研究者

本文提出一种多模态架构,用于从战术摄像机画面中分析足球场景,聚焦三个核心任务:球轨迹推断、球状态分类和持球者识别。该方法整合三种输入模态(球员轨迹、球员类型、个体球员图像块),利用一系列社会时序变换器块处理时空动态。与以往依赖精确球跟踪或手工启发式规则的方法不同,本方法在无法获取球过去或未来位置的情况下推断球轨迹,并在真实顶级联赛比赛中存在噪声或遮挡时,仍能稳健识别球状态与持球者。我们引入CropDrop——一种模态特异性遮蔽预训练策略,防止模型过度依赖图像特征,促使模型在预训练阶段更关注跨模态模式。在大规模数据集上的实验表明,本方法在所有任务中均显著优于现有最先进基线。结果凸显了基于变换器架构中结构化与视觉线索结合的优势,以及真实遮蔽策略在多模态学习中的重要性。

原文摘要 · Abstract (English)

In this work we propose a multi-modal architecture for analyzing soccer scenes from tactical camera footage, with a focus on three core tasks: ball trajectory inference, ball state classification, and ball possessor identification. To this end, our solution integrates three distinct input modalities (player trajectories, player types and image crops of individual players) into a unified framework that processes spatial and temporal dynamics using a cascade of sociotemporal transformer blocks. Unlike prior methods, which rely heavily on accurate ball tracking or handcrafted heuristics, our approach infers the ball trajectory without direct access to its past or future positions, and robustly identifies the ball state and ball possessor under noisy or occluded conditions from real top league matches. We also introduce CropDrop, a modality-specific masking pre-training strategy that prevents over-reliance on image features and encourages the model to rely on cross-modal patterns during pre-training. We show the effectiveness of our approach on a large-scale dataset providing substantial improvements over state-of-the-art baselines in all tasks. Our results highlight the benefits of combining structured and visual cues in a transformer-based architecture, and the importance of realistic masking strategies in multi-modal learning.

足球分析多模态变换器轨迹推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。