arXiv:2508.11569cs.CVcs.IR2025-08中稿 · TCSVT被引 1

TrajSV用轨迹建模提升体育视频理解,无监督训练效果显著。

TrajSV: A Trajectory-based Model for Sports Video Representations and Applications

  • 基于球员与球的轨迹构建视频表征框架
  • 在3个数据集上实现近70%检索性能提升
  • 适合体育视频分析、动作定位等应用开发

近年来,体育数据分析受到学术界和产业界的广泛关注。尽管研究不断深入,仍存在数据缺失、缺乏有效的轨迹建模框架以及依赖大量标注等问题。本文提出TrajSV,一种基于轨迹的体育视频表征框架,包含数据预处理、片段表示网络(CRNet)和视频表示网络(VRNet)三部分。预处理模块从体育转播视频中提取球员与球的轨迹;CRNet利用轨迹增强的Transformer学习片段表征;VRNet通过编码器-解码器结构聚合片段表征与视觉特征,生成视频表征。引入三重对比损失,在无监督条件下优化表征。在三个转播视频数据集上验证了TrajSV在足球、篮球、排球三种运动中的有效性,涵盖视频检索、动作定位和视频描述三个下游任务。实验结果表明,该方法在体育视频检索中达到领先水平,性能提升近70%;在动作定位中,17个动作类别中有9个达到最优,表现超越基线;在视频描述任务中实现近20%的提升。此外,本文还部署了基于TrajSV的系统及三项应用。

原文摘要 · Abstract (English)

Sports analytics has received significant attention from both academia and industry in recent years. Despite the growing interest and efforts in this field, several issues remain unresolved, including (1) data unavailability, (2) lack of an effective trajectory-based framework, and (3) requirement for sufficient supervision labels. In this paper, we present TrajSV, a trajectory-based framework that addresses various issues in existing studies. TrajSV comprises three components: data preprocessing, Clip Representation Network (CRNet), and Video Representation Network (VRNet). The data preprocessing module extracts player and ball trajectories from sports broadcast videos. CRNet utilizes a trajectory-enhanced Transformer module to learn clip representations based on these trajectories. Additionally, VRNet learns video representations by aggregating clip representations and visual features with an encoder-decoder architecture. Finally, a triple contrastive loss is introduced to optimize both video and clip representations in an unsupervised manner. The experiments are conducted on three broadcast video datasets to verify the effectiveness of TrajSV for three types of sports (i.e., soccer, basketball, and volleyball) with three downstream applications (i.e., sports video retrieval, action spotting, and video captioning). The results demonstrate that TrajSV achieves state-of-the-art performance in sports video retrieval, showcasing a nearly 70% improvement. It outperforms baselines in action spotting, achieving state-of-the-art results in 9 out of 17 action categories, and demonstrates a nearly 20% improvement in video captioning. Additionally, we introduce a deployed system along with the three applications based on TrajSV.

体育视频轨迹建模无监督学习视频理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。