arXiv:2501.07221cs.CVcs.AI2025-01被引 8

用CLIP模型实现高精度瑜伽体位识别,支持实时应用。

Exploring the Use of Contrastive Language-Image Pre-Training for Human Posture Classification: Insights from Yoga Pose Analysis

  • 基于CLIP进行迁移学习,优化图像描述语法与超参数
  • 在82类数据上达85%准确率,比现有方法高6%,训练时间仅为YOLOv8的1/3.5
  • 小样本下(20张/类)仍可达90%准确率,适合实际部署

准确的人体姿态分类对工作安全、康复训练、体育教学及日常生活辅助等场景至关重要。近年来,对比语言-图像预训练(CLIP)在图文联合理解方面取得显著进展。本研究评估了CLIP在人体姿态分类中的有效性,聚焦瑜伽体位识别。尽管零样本方法存在局限,但通过对15,301张图像(真实与合成)进行微调,涵盖82个类别,获得良好效果。文章详细描述了微调流程,包括图像描述语法选择、模型与超参数调整。微调后的CLIP模型在3,826张测试图像上达到超过85%的准确率,较同类数据集上现有方法提升约6%,且训练时间仅为微调YOLOv8模型的1/3.5。在更贴近应用的小样本场景中,六类体位分别使用1,301和401张训练图像,准确率分别达到98.8%和99.1%。实验还表明,每类仅需20张图像即可实现约90%准确率。结果表明,该多模态方法可有效应用于瑜伽体位识别,乃至一般人体姿态分类。此外,CLIP推理时间约为7毫秒,支持集成至自动化系统,如开发实时个人瑜伽助教用于动作评估。

原文摘要 · Abstract (English)

Accurate human posture classification in images and videos is crucial for automated applications across various fields, including work safety, physical rehabilitation, sports training, or daily assisted living. Recently, multimodal learning methods, such as Contrastive Language-Image Pretraining (CLIP), have advanced significantly in jointly understanding images and text. This study aims to assess the effectiveness of CLIP in classifying human postures, focusing on its application in yoga. Despite the initial limitations of the zero-shot approach, applying transfer learning on 15,301 images (real and synthetic) with 82 classes has shown promising results. The article describes the full procedure for fine-tuning, including the choice for image description syntax, models and hyperparameters adjustment. The fine-tuned CLIP model, tested on 3826 images, achieves an accuracy of over 85%, surpassing the current state-of-the-art of previous works on the same dataset by approximately 6%, its training time being 3.5 times lower than what is needed to fine-tune a YOLOv8-based model. For more application-oriented scenarios, with smaller datasets of six postures each, containing 1301 and 401 training images, the fine-tuned models attain an accuracy of 98.8% and 99.1%, respectively. Furthermore, our experiments indicate that training with as few as 20 images per pose can yield around 90% accuracy in a six-class dataset. This study demonstrates that this multimodal technique can be effectively used for yoga pose classification, and possibly for human posture classification, in general. Additionally, CLIP inference time (around 7 ms) supports that the model can be integrated into automated systems for posture evaluation, e.g., for developing a real-time personal yoga assistant for performance assessment.

姿态识别CLIP小样本实时系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。