用视频级标注训练实时评估中风康复动作的模型,减少人工标注负担。
Frame-Level Real-Time Assessment of Stroke Rehabilitation Exercises from Video-Level Labeled Data: Task-Specific vs. Foundation Models
- 通过梯度方法和伪标签选择,从视频级标签生成帧级伪标签。
- 基于预训练模型的框架在帧级评估上达到AUC 72%,优于基线69%。
- 适用于新患者快速适配,降低康复动作分析的数据标注成本。
中风康复需求增长推动了自主锻炼支持系统的发展。虚拟教练可基于视频数据提供实时动作反馈,帮助患者改善运动功能并保持参与度。然而,训练实时动作分析系统需帧级标注,耗时且昂贵。本文提出一种框架,仅使用视频级标注即可学习分类单帧,用于康复动作中代偿行为的实时评估。采用基于梯度的技术与伪标签选择方法生成帧级伪标签,并利用预训练任务特定模型(Action Transformer、SkateFormer)及基础模型(MOMENT)生成伪标签,以提升对新患者的泛化能力。在包含18名中风患者执行五种康复动作的SERE数据集上验证:MOMENT在视频级评估中取得AUC 73%,优于基线LSTM的58%;Action Transformer结合积分梯度法在帧级评估中达AUC 72%,超过使用真实帧级标签训练的基线(AUC 69%)。结果表明,该方法能增强模型泛化性,实现对新患者的快速定制,显著降低标注需求。
原文摘要 · Abstract (English)
The growing demands of stroke rehabilitation have increased the need for solutions to support autonomous exercising. Virtual coaches can provide real-time exercise feedback from video data, helping patients improve motor function and keep engagement. However, training real-time motion analysis systems demands frame-level annotations, which are time-consuming and costly to obtain. In this work, we present a framework that learns to classify individual frames from video-level annotations for real-time assessment of compensatory motions in rehabilitation exercises. We use a gradient-based technique and a pseudo-label selection method to create frame-level pseudo-labels for training a frame-level classifier. We leverage pre-trained task-specific models - Action Transformer, SkateFormer - and a foundation model - MOMENT - for pseudo-label generation, aiming to improve generalization to new patients. To validate the approach, we use the \textit{SERE} dataset with 18 post-stroke patients performing five rehabilitation exercises annotated on compensatory motions. MOMENT achieves better video-level assessment results (AUC = $73\%$), outperforming the baseline LSTM (AUC = $58\%$). The Action Transformer, with the Integrated Gradient technique, leads to better outcomes (AUC = $72\%$) for frame-level assessment, outperforming the baseline trained with ground truth frame-level labeling (AUC = $69\%$). We show that our proposed approach with pre-trained models enhances model generalization ability and facilitates the customization to new patients, reducing the demands of data labeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。