arXiv:2501.03674cs.CVcs.AI2025-01被引 31

通过分层姿态引导的多阶段对比回归,提升动作质量评估精度。

Action Quality Assessment via Hierarchical Pose-guided Multi-stage Contrastive Regression

  • 分层姿态引导,捕捉细微动作差异
  • 多阶段对比回归,准确预测动作质量分数
  • 新标注数据集支持,适合运动分析研究者

动作质量评估(AQA)旨在自动、公正地评价运动员表现,近年来受到广泛关注。然而,运动员常处于快速运动中,视觉外观差异微小,难以捕捉精细姿态变化,导致评估性能不佳。此外,多数AQA任务(如跳水)包含多个持续时间不同的子动作,现有方法通常将视频划分为固定帧,破坏了子动作的时间连续性,引发预测误差。为此,我们提出一种基于分层姿态引导的多阶段对比回归方法。首先,设计多尺度动态视觉-骨骼编码器,捕获细粒度时空特征;其次,引入过程分割网络分离不同子动作并获取分段特征;随后,将分段的视觉与骨骼特征输入多模态融合模块,作为物理结构先验,指导模型学习精确的动作相似性与差异性;最后,采用多阶段对比学习回归策略,生成判别性表征并输出预测结果。此外,我们构建了新标注的FineDiving-Pose数据集,以改善现有低质量人体姿态标签问题。在FineDiving和MTL-AQA数据集上的实验表明,所提方法有效且具有优势。源代码与数据集已公开于https://github.com/Lumos0507/HP-MCoRe。

原文摘要 · Abstract (English)

Action Quality Assessment (AQA), which aims at automatic and fair evaluation of athletic performance, has gained increasing attention in recent years. However, athletes are often in rapid movement and the corresponding visual appearance variances are subtle, making it challenging to capture fine-grained pose differences and leading to poor estimation performance. Furthermore, most common AQA tasks, such as diving in sports, are usually divided into multiple sub-actions, each of which contains different durations. However, existing methods focus on segmenting the video into fixed frames, which disrupts the temporal continuity of sub-actions resulting in unavoidable prediction errors. To address these challenges, we propose a novel action quality assessment method through hierarchically pose-guided multi-stage contrastive regression. Firstly, we introduce a multi-scale dynamic visual-skeleton encoder to capture fine-grained spatio-temporal visual and skeletal features. Then, a procedure segmentation network is introduced to separate different sub-actions and obtain segmented features. Afterwards, the segmented visual and skeletal features are both fed into a multi-modal fusion module as physics structural priors, to guide the model in learning refined activity similarities and variances. Finally, a multi-stage contrastive learning regression approach is employed to learn discriminative representations and output prediction results. In addition, we introduce a newly-annotated FineDiving-Pose Dataset to improve the current low-quality human pose labels. In experiments, the results on FineDiving and MTL-AQA datasets demonstrate the effectiveness and superiority of our proposed approach. Our source code and dataset are available at https://github.com/Lumos0507/HP-MCoRe.

动作评估姿态识别多阶段学习对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。