arXiv:2508.10281cs.CV2025-08

提出新方法精准识别花样滑冰跳跃动作的类型与时机。

VIFSS: View-Invariant and Figure Skating-Specific Pose Representation Learning for Temporal Action Segmentation

  • 构建3D姿态表示学习框架,实现视角不变性
  • 在有限数据下达成92%以上跳动识别准确率
  • 适合体育动作分析与小样本场景应用

从视频中理解人类动作在多个领域至关重要,尤其在体育分析中。在花样滑冰中,准确识别运动员完成的跳跃类型与时间对客观评分极为关键,但因动作细粒度高、结构复杂,通常需专家知识。现有时间动作分割(TAS)方法面临两大挑战:标注数据不足,且未考虑跳跃动作固有的三维特性与流程结构。本文提出针对花样滑冰跳跃的时间动作分割新框架(VIFSS),显式融合三维特征与语义流程结构。首先,提出一种视图不变、花样滑冰专用的姿态表示学习方法,采用对比学习预训练与动作分类微调相结合;为此构建了首个公开的3D姿态数据集FS-Jump3D。其次,设计细粒度标注方案,标记“起始(准备)”和“落地”阶段,使模型学习跳跃流程。大量实验表明,该方法在元素级TAS任务中达到92%以上的F1@50,能同时识别跳跃类型与旋转等级。此外,在微调数据稀缺时,视图不变对比预训练仍表现优异,凸显其在真实场景中的实用性。

原文摘要 · Abstract (English)

Understanding human actions from videos plays a critical role across various domains, including sports analytics. In figure skating, accurately recognizing the type and timing of jumps a skater performs is essential for objective performance evaluation. However, this task typically requires expert-level knowledge due to the fine-grained and complex nature of jump procedures. While recent approaches have attempted to automate this task using Temporal Action Segmentation (TAS), there are two major limitations to TAS for figure skating: the annotated data is insufficient, and existing methods do not account for the inherent three-dimensional aspects and procedural structure of jump actions. In this work, we propose a new TAS framework for figure skating jumps that explicitly incorporates both the three-dimensional nature and the semantic procedure of jump movements. First, we propose a novel View-Invariant, Figure Skating-Specific pose representation learning approach (VIFSS) that combines contrastive learning as pre-training and action classification as fine-tuning. For view-invariant contrastive pre-training, we construct FS-Jump3D, the first publicly available 3D pose dataset specialized for figure skating jumps. Second, we introduce a fine-grained annotation scheme that marks the ``entry (preparation)'' and ``landing'' phases, enabling TAS models to learn the procedural structure of jumps. Extensive experiments demonstrate the effectiveness of our framework. Our method achieves over 92% F1@50 on element-level TAS, which requires recognizing both jump types and rotation levels. Furthermore, we show that view-invariant contrastive pre-training is particularly effective when fine-tuning data is limited, highlighting the practicality of our approach in real-world scenarios.

动作分割3D姿态花滑分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。