arXiv:2503.00147cs.CV2025-03CVPR被引 7

提出新模型精准定位体育视频中的事件,解决长距离依赖与类别不平衡问题。

Precise Event Spotting in Sports Videos: Solving Long-Range Dependency and Class Imbalance

  • 设计融合时空精修与长程时序建模的网络结构。
  • 在挑战性场景下显著优于当前最佳方法,尤其提升稀有事件识别率。
  • 适合需要高精度事件定位的体育视频分析任务。

精确事件定位(PES)旨在从长且未剪辑的体育视频中准确识别事件及其类别,核心目标是精确捕捉事件发生的时刻。现有方法主要依赖大型预训练网络提取特征,但未必适合该任务;同时忽略了数据中存在的类别分布不均问题,影响了复杂场景下的性能。本文表明,经合理设计并端到端训练的网络可超越当前最优方法。我们提出一种结合卷积时空特征提取器与自适应时空精修模块(ASTRM)及长程时序模块的网络架构。ASTRM增强时空特征表达,长程时序模块则通过建模长距离依赖关系捕捉全局上下文。为应对类别不平衡,引入软实例对比损失(SoftIC),促进特征紧凑性和类别分离。大量实验表明,所提方法高效且在更具挑战性的设置下优于现有SOTA方法。

原文摘要 · Abstract (English)

Precise Event Spotting (PES) aims to identify events and their class from long, untrimmed videos, particularly in sports. The main objective of PES is to detect the event at the exact moment it occurs. Existing methods mainly rely on features from a large pre-trained network, which may not be ideal for the task. Furthermore, these methods overlook the issue of imbalanced event class distribution present in the data, negatively impacting performance in challenging scenarios. This paper demonstrates that an appropriately designed network, trained end-to-end, can outperform state-of-the-art (SOTA) methods. Particularly, we propose a network with a convolutional spatial-temporal feature extractor enhanced with our proposed Adaptive Spatio-Temporal Refinement Module (ASTRM) and a long-range temporal module. The ASTRM enhances the features with spatio-temporal information. Meanwhile, the long-range temporal module helps extract global context from the data by modeling long-range dependencies. To address the class imbalance issue, we introduce the Soft Instance Contrastive (SoftIC) loss that promotes feature compactness and class separation. Extensive experiments show that the proposed method is efficient and outperforms the SOTA methods, specifically in more challenging settings.

事件定位体育视频长程依赖类别平衡

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。