arXiv:2608.01696cs.CVcs.AI2026-08

提出新模型精准定位球员持球动作并区分角色。

Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting

论文配图:Entity-Aware Sequence Transduction for Player-Centric Ball Action Spotting
图 1 · 摘自论文原文
  • 保留球员角色维度,分步建模个体演化与互动。
  • 在FOOTPASS数据集上微平均F1达0.778,提升10.3个百分点。
  • 适合需要精确球员行为分析的体育视频研究者。

以球员为中心的持球动作检测需在拥挤、部分观测的多智能体体育视频中实现高时序精度的动作识别与主体归属。现有去噪序列转换(DST)基线将球员角色维度融合进帧级表示,削弱了对球员特异性时序演化及交互建模的归纳偏置。为此,本文提出多实体去噪序列转换(ME-DST),在编码过程中保持角色槽维度。通过时间注意力建模每个角色槽的历史,空间注意力在每帧跨角色槽传递信息。这种分解设计使模型能直接分离个体演化与角色间上下文。此外,引入可学习角色嵌入、追踪生成的战术特征,以及来自X3D-L和Swin3D-S的融合视觉预测。在FOOTPASS数据集上的实验表明,ME-DST取得0.778的微平均F1,较最强官方基线TAAD+DST提升10.3个百分点。控制消融实验显示,保持实体轴和编码角色身份是性能提升的关键。结果表明,显式实体建模是球员中心体育事件理解的有效归纳偏置。

原文摘要 · Abstract (English)

Player-centric ball action spotting requires temporally precise event detection together with actor attribution in crowded, partially observed multi-agent sports videos. Existing Denoising Sequence Transduction (DST) baselines treat the player-role dimension as part of a flattened frame-level representation, which weakens the inductive bias for modeling player-specific temporal evolution and inter-player interactions. To address this limitation, we propose Multi-Entity Denoising Sequence Transduction (ME-DST). ME-DST keeps the role-slot dimension throughout encoding. It uses temporal attention to model the history of each role slot, and spatial attention to exchange information across role slots at each frame. This factorized design gives the model a direct structure for separating within-player evolution from inter-player context. We also add learnable role embeddings, tracking-derived tactical features, and fused visual predictions from X3D-L and Swin3D-S. Experiments on the FOOTPASS dataset show that ME-DST reaches a Micro F1 of 0.778. This improves the strongest official TAAD+DST baseline by 10.3 percentage points. Controlled ablations show that preserving the entity axis and encoding role identity are central to this gain. These results suggest that explicit entity modeling is an effective inductive bias for player-centric sports event understanding.

动作识别体育分析序列建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。