arXiv:2511.17727cs.CV2025-11被引 3

用视觉语言模型分析中风康复视频,发现精度有限但有改进潜力。

The Potential and Limitations of Vision-Language Models for Human Motion Understanding: A Case Study in Data-Driven Stroke Rehabilitation

  • 将康复动作识别建模为视觉-语言任务,不需微调直接推理。
  • 轻度受损者剂量估计误差小于25%,但严重患者无法准确预测。
  • 无需训练即可识别高阶动作和基本运动,适合临床视频初筛。

视觉语言模型(VLMs)在众多计算机视觉任务中表现优异,引发其在数字健康领域的应用兴趣。本文将其应用于数据驱动中风康复中的两大核心挑战:康复剂量与功能障碍的自动量化。我们将问题建模为运动识别任务,利用VLMs进行求解。在包含29名健康对照与51名中风患者的队列上评估该框架。结果表明,当前VLMs缺乏精细运动理解能力,剂量估计与不使用视觉信息的基线相当,功能障碍评分无法可靠预测。然而,优化提示与后处理后,模型可从少量帧中分类高阶活动,以中等精度检测运动与抓握行为,并对轻度受损及健康参与者实现剂量计数误差低于25%的近似值,且无需任务特定训练或微调。这些结果揭示了VLMs在中风康复及更广泛临床视频分析中的当前局限与潜在机遇。

原文摘要 · Abstract (English)

Vision-language models (VLMs) have demonstrated remarkable performance across a wide range of computer-vision tasks, sparking interest in their potential for digital health applications. Here, we apply VLMs to two fundamental challenges in data-driven stroke rehabilitation: automatic quantification of rehabilitation dose and impairment from videos. We formulate these problems as motion-identification tasks, which can be addressed using VLMs. We evaluate our proposed framework on a cohort of 29 healthy controls and 51 stroke survivors. Our results show that current VLMs lack the fine-grained motion understanding required for precise quantification: dose estimates are comparable to a baseline that excludes visual information, and impairment scores cannot be reliably predicted. Nevertheless, several findings suggest future promise. With optimized prompting and post-processing, VLMs can classify high-level activities from a few frames, detect motion and grasp with moderate accuracy, and approximate dose counts within 25% of ground truth for mildly impaired and healthy participants, all without task-specific training or finetuning. These results highlight both the current limitations and emerging opportunities of VLMs for data-driven stroke rehabilitation and broader clinical video analysis.

中风康复视觉语言模型动作识别临床视频分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。