提出多焦点时移模块,精准定位体育视频中的细微动作。
Multi-Focus Temporal Shifting for Precise Event Spotting in Sports Videos
- 用多尺度时移与分组聚焦机制增强时序建模能力。
- 在五个数据集上提升4.09 mAP,仅需45 GFLOPs计算量。
- 适合需要轻量化高精度动作识别的体育视频场景。
体育视频中的精确事件定位(PES)要求从单摄像头画面中进行帧级细粒度动作识别。现有PES模型通常使用轻量级时序模块(如门控时移模块GSM)来为2D CNN特征提取器补充时序上下文,但这些模块在时序感受野和空间适应性方面存在局限。本文提出多焦点时移模块(MFS),通过引入多尺度时移和分组聚焦机制,在保持高效的同时,有效建模短时与长时依赖关系,并聚焦显著区域。MFS为轻量级、即插即用模块,可无缝集成于多种2D骨干网络。为进一步推动该领域发展,我们构建了首个乒乓球事件定位基准数据集Table Tennis Australia,包含超过4,800个精标注事件。在五个PES基准上的大量实验表明,MFS在极少额外开销下持续提升性能,成为轻量级方法中的领先者(+4.09 mAP,45 GFLOPs)。
原文摘要 · Abstract (English)
Precise Event Spotting (PES) in sports videos requires frame-level recognition of fine-grained actions from single-camera footage. Existing PES models typically incorporate lightweight temporal modules such as the Gate Shift Module (GSM) or the Gate Shift Fuse to enrich 2D CNN feature extractors with temporal context. However, these modules are limited in both temporal receptive field and spatial adaptability. We propose Multi-Focus Temporal Shifting Module (MFS) that enhances GSM with multi-scale temporal shifts and Group Focus Module, enabling efficient modeling of both short and long-term dependencies while focusing on salient regions. MFS is a lightweight, plug-and-play module that integrates seamlessly with diverse 2D backbones. To further advance the field, we introduce the Table Tennis Australia dataset, the first PES benchmark for table tennis containing over 4,800 precisely annotated events. Extensive experiments across five PES benchmarks demonstrate that MFS consistently improves performance with minimal overhead, achieving leading results among lightweight methods (+4.09 mAP, 45 GFLOPs).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。