arXiv:2512.00572cs.CVcs.AI2025-12被引 2

用骨骼特征提升瑜伽动作识别准确率,最高达96.09%

Integrating Skeleton Based Representations for Robust Yoga Pose Classification Using Deep Learning Models

  • 采用骨骼关键点作为输入,比直接使用图像更有效
  • 基于MediaPipe的骨骼数据使VGG16达到96.09%准确率
  • 适合关注动作识别与模型可解释性的研究者

瑜伽因其身心益处广受欢迎,但错误姿势易致伤,自动化姿态识别因而重要。尽管人体姿态关键点提取模型在动作识别中表现优异,现有研究对瑜伽姿态识别的系统性评估仍不足,多仅依赖原始图像或单一姿态提取模型。本研究提出新数据集Yoga-16,克服现有数据集局限,并系统评估VGG16、ResNet50和Xception三种深度学习架构,使用三种输入模态:原始图像、MediaPipe Pose骨骼图、YOLOv8 Pose骨骼图。实验表明,骨骼表示优于原始图像输入,其中VGG16结合MediaPipe Pose骨骼输入取得96.09%最高准确率。同时通过Grad-CAM进行可解释性分析,揭示模型决策依据,并完成交叉验证。

原文摘要 · Abstract (English)

Yoga is a popular form of exercise worldwide due to its spiritual and physical health benefits, but incorrect postures can lead to injuries. Automated yoga pose classification has therefore gained importance to reduce reliance on expert practitioners. While human pose keypoint extraction models have shown high potential in action recognition, systematic benchmarking for yoga pose recognition remains limited, as prior works often focus solely on raw images or a single pose extraction model. In this study, we introduce a curated dataset, 'Yoga-16', which addresses limitations of existing datasets, and systematically evaluate three deep learning architectures (VGG16, ResNet50, and Xception), using three input modalities (direct images, MediaPipe Pose skeleton images, and YOLOv8 Pose skeleton images). Our experiments demonstrate that skeleton-based representations outperform raw image inputs, with the highest accuracy of 96.09% achieved by VGG16 with MediaPipe Pose skeleton input. Additionally, we provide interpretability analysis using Grad-CAM, offering insights into model decision-making for yoga pose classification with cross-validation analysis.

姿态识别骨骼特征深度学习瑜伽

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。