用视觉变压器检测美式橄榄球训练中的危险冲撞,提升安全预警能力。
ViTs for Action Classification in Videos: An Approach to Risky Tackle Detection in American Football Practice Videos
- 基于视觉变压器模型,结合不平衡数据处理策略
- 在大规模数据上实现0.67的危险动作召回率和0.59的F1值
- 适合教练团队用于球员受伤预防,尤其关注罕见高危动作
接触类运动中早期识别危险动作可及时干预,提升运动员安全。本文提出一种检测美式橄榄球训练视频中危险冲撞的方法,并构建了一个大幅扩展的数据集。该数据集包含733个单人与假人冲撞片段,每个片段均以首次接触时刻为中心进行时间定位,并依据标准化冲撞技术评估量表(SATT-3)标注击打区域,较之前工作增加超过四倍(原为178段)。采用基于视觉变压器的模型并引入类别不平衡感知训练,在交叉验证下取得0.67的危险召回率和0.59的危险F1值。相较此前在小规模子集上的基线表现(危险召回率0.58,F1 0.56),本方法在更大数据集上提升超过8个百分点。结果表明,结合视觉变压器与精细不平衡处理的视频分析,可有效识别稀有但关键的安全隐患动作,为教练主导的伤防工具提供可行路径。
原文摘要 · Abstract (English)
Early identification of hazardous actions in contact sports enables timely intervention and improves player safety. We present a method for detecting risky tackles in American football practice videos and introduce a substantially expanded dataset for this task. Our work contains 733 single-athlete-dummy tackle clips, each temporally localized around first point contact and labeled with a strike zone component of the standardized Assessment for Tackling Technique (SATT-3), extending prior work that reported 178 annotated videos. Using a Vision transformer-based model with imbalance-aware training, we obtain risky recall of 0.67 and Risky F1 of 0.59 under crossvalidation. Relative to the previous baseline in a smaller subset (risky recall of 0.58; Risky F1 0.56 ), our approach improves risky recall by more than 8% points on a much larger dataset. These results indicate that the vision transformer-based video analysis, coupled with careful handling of class imbalance, can reliably detect rare but safety-critical tackling patterns, offering a practical pathway toward coach-centered injury prevention tools.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。