对比视频与位置数据在团队活动识别中的效果,发现位置信息更优。
Pixels or Positions? Benchmarking Modalities in Group Activity Recognition
- 用足球世界杯数据构建同步视频与追踪数据集
- 基于位置的模型准确率达77.8%,远超视频模型的60.9%
- 仅需少量参数和计算量,适合高效部署
群体活动识别(GAR)在监控与室内团队运动(如排球、篮球)中研究广泛,但轨迹等位置信号仍较少被探索。本文引入SoccerNet-GAR数据集,基于2022年世界杯64场比赛的广播视频与球员追踪数据,对87,939个群体活动进行标注,涵盖10类动作。我们建立统一评估协议,对比视频与追踪两类方法:前者使用强基线分类器,后者采用新型角色感知图神经网络,通过球员场上角色连接位置边来建模战术结构。结果显示,追踪模型达到77.8%平衡准确率,优于最佳视频基线的60.9%;训练仅需7倍少的GPU小时与479倍少的参数(18万 vs 8630万),显著提升效率。
原文摘要 · Abstract (English)
Group Activity Recognition (GAR) is well studied on the video modality for surveillance and indoor team sports (e.g., volleyball, basketball). Yet, other modalities such as agent positions and trajectories over time, i.e. tracking, remain comparatively under-explored despite being compact, agent-centric signals that explicitly encode spatial interactions. Understanding whether pixel (video) or position (tracking) modalities leads to better group activity recognition is therefore important to drive further research on the topic. However, no standardized benchmark currently exists that aligns broadcast video and tracking data for the same group activities, leading to a lack of apples-to-apples comparison between these modalities for GAR. In this work, we introduce SoccerNet-GAR, a multimodal dataset built from the $64$ matches of the football World Cup 2022. Specifically, the broadcast videos and player tracking modalities for $87{,}939$ group activities are synchronized and annotated with $10$ categories. Furthermore, we define a unified evaluation protocol to benchmark two strong unimodal approaches: (i) competitive video-based classifiers and (ii) tracking-based classifiers leveraging graph neural networks. In particular, our novel role-aware graph architecture for tracking-based GAR directly encodes tactical structure through positional edges connecting players by their on-pitch roles. Our tracking model achieves $77.8\%$ balanced accuracy compared to $60.9\%$ for the best video baseline, while training with $7 \times$ less GPU hours and $479 \times$ fewer parameters ($180K$ vs. $86.3M$). This study provides new insights into the relative strengths of pixels and positions for group activity recognition in sports.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。