首个基于3D骨骼的团队活动理解基准,聚焦篮球战术复杂交互
SGA-INTERACT: A 3D Skeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic
- 构建首个3D骨骼数据集,支持长时序和多视角活动分析
- 提出时间组活动定位任务,覆盖非剪切视频序列
- 设计统一特征提取框架,兼容RGB与骨骼模型评估
群体活动理解主要以群体活动识别(GAR)任务为主,但现有GAR基准存在活动类别粗粒度、仅支持单视角数据的问题,限制了先进算法的评估。为此,我们提出SGA-INTERACT,首个基于3D骨骼的群体活动理解基准,包含受篮球战术启发的复杂活动,强调丰富的空间交互与长时依赖。该基准引入时间组活动定位(TGAL)任务,将群体活动理解扩展至未剪切序列,填补了GAR作为独立任务的空白。此外,我们提出One2Many框架,采用预训练3D骨骼主干网络实现统一个体特征提取,与基于RGB的方法特征提取范式对齐,支持直接在骨骼基准上评估RGB模型。我们在SGA-INTERACT上对两种骨骼方法、三种RGB方法及一个基于One2Many的基线进行了广泛评估,基线普遍表现不佳,凸显该基准的挑战性,推动群体活动理解的发展。
原文摘要 · Abstract (English)
Group Activity Understanding is predominantly studied as Group Activity Recognition (GAR) task. However, existing GAR benchmarks suffer from coarse-grained activity vocabularies and the only data form in single-view, which hinder the evaluation of state-of-the-art algorithms. To address these limitations, we introduce SGA-INTERACT, the first 3D skeleton-based benchmark for group activity understanding. It features complex activities inspired by basketball tactics, emphasizing rich spatial interactions and long-term dependencies. SGA-INTERACT introduces Temporal Group Activity Localization (TGAL) task, extending group activity understanding to untrimmed sequences, filling the gap left by GAR as a standalone task. In addition to the benchmark, we propose One2Many, a novel framework that employs a pretrained 3D skeleton backbone for unified individual feature extraction. This framework aligns with the feature extraction paradigm in RGB-based methods, enabling direct evaluation of RGB-based models on skeleton-based benchmarks. We conduct extensive evaluations on SGA-INTERACT using two skeleton-based methods, three RGB-based methods, and a proposed baseline within the One2Many framework. The general low performance of baselines highlights the benchmark's challenges, motivating advancements in group activity understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。