用手机摄像头实现上百种运动的远程监测,精度高且部署简单。
(LiFT) Lightweight Fitness Transformer: A language-vision model for Remote Monitoring of Physical Training
- 构建视觉语言模型,统一处理动作识别与重复计数。
- 在1900+动作数据集上,识别准确率76.5%,计数误差≤1次达85.3%。
- 仅需普通手机摄像头,适合大众化健身追踪应用。
我们提出一种仅需RGB智能手机摄像头的健身追踪系统,实现远程健身监控,更具隐私性、可扩展性和低成本优势。尽管已有自动化锻炼监督研究,但现有模型要么动作种类有限,要么过于复杂难以实际部署。以往方法通常局限于少数动作,难以泛化到多样化运动。相比之下,我们开发了一种稳健的多任务运动分析模型,可识别并计数数百种动作,远超此前方法。通过构建大规模健身数据集Olympia(含超过1,900种动作),克服了先前的数据局限。据我们所知,该视觉-语言模型是首个能对骨骼运动数据执行多项任务的模型。在Olympia数据集上,模型在仅使用RGB视频的情况下,实现了76.5%的动作检测准确率和85.3%的离群值≤1的重复计数准确率。通过单一视觉-语言变压器模型完成动作识别与重复计数,我们朝着实现普惠型AI健身追踪迈出了重要一步。
原文摘要 · Abstract (English)
We introduce a fitness tracking system that enables remote monitoring for exercises using only a RGB smartphone camera, making fitness tracking more private, scalable, and cost effective. Although prior work explored automated exercise supervision, existing models are either too limited in exercise variety or too complex for real-world deployment. Prior approaches typically focus on a small set of exercises and fail to generalize across diverse movements. In contrast, we develop a robust, multitask motion analysis model capable of performing exercise detection and repetition counting across hundreds of exercises, a scale far beyond previous methods. We overcome previous data limitations by assembling a large-scale fitness dataset, Olympia covering more than 1,900 exercises. To our knowledge, our vision-language model is the first that can perform multiple tasks on skeletal fitness data. On Olympia, our model can detect exercises with 76.5% accuracy and count repetitions with 85.3% off-by-one accuracy, using only RGB video. By presenting a single vision-language transformer model for both exercise identification and rep counting, we take a significant step toward democratizing AI-powered fitness tracking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。