首个专用于中风康复的多模态动作数据集,助力客观评估上肢恢复进展。
STROKEVISION-BENCH: A Multimodal Video And 2D Pose Benchmark For Tracking Stroke Recovery
- 构建中风患者完成标准化方块转移任务的视频与2D骨骼关键点数据集
- 包含1000段标注视频,分4类临床动作,支持多模态分析
- 为自动化康复评估提供基准,适合康复医学与计算机视觉交叉研究
尽管康复方案不断进步,中风后上肢功能的临床评估仍以治疗师观察和粗略评分为主,主观性强,难以捕捉细微运动改善,影响个性化康复计划制定。近年来计算机视觉技术为实现客观、量化、可扩展的上肢功能评估提供了新路径。在标准化测试中,方块传递测试(Box and Block Test, BBT)被广泛用于评估手部粗大协调能力并追踪中风恢复进程,其结构化场景适合计算分析。然而,现有康复数据集多聚焦日常生活活动,缺乏如方块转移等临床结构化任务的记录;且多数数据混合健康人与中风患者,限制了临床针对性。为此,我们提出StrokeVision-Bench——首个专为中风患者设计的结构化方块转移任务数据集。该数据集包含1000段标注视频,分为4类临床有意义的动作类别,每条样本均以原始视频帧和2D骨骼关键点两种模态呈现。我们对多种先进的视频动作识别与基于骨架的动作分类方法进行基准测试,建立了该领域的性能基线,推动自动化中风康复评估研究发展。
原文摘要 · Abstract (English)
Despite advancements in rehabilitation protocols, clinical assessment of upper extremity (UE) function after stroke largely remains subjective, relying heavily on therapist observation and coarse scoring systems. This subjectivity limits the sensitivity of assessments to detect subtle motor improvements, which are critical for personalized rehabilitation planning. Recent progress in computer vision offers promising avenues for enabling objective, quantitative, and scalable assessment of UE motor function. Among standardized tests, the Box and Block Test (BBT) is widely utilized for measuring gross manual dexterity and tracking stroke recovery, providing a structured setting that lends itself well to computational analysis. However, existing datasets targeting stroke rehabilitation primarily focus on daily living activities and often fail to capture clinically structured assessments such as block transfer tasks. Furthermore, many available datasets include a mixture of healthy and stroke-affected individuals, limiting their specificity and clinical utility. To address these critical gaps, we introduce StrokeVision-Bench, the first-ever dedicated dataset of stroke patients performing clinically structured block transfer tasks. StrokeVision-Bench comprises 1,000 annotated videos categorized into four clinically meaningful action classes, with each sample represented in two modalities: raw video frames and 2D skeletal keypoints. We benchmark several state-of-the-art video action recognition and skeleton-based action classification methods to establish performance baselines for this domain and facilitate future research in automated stroke rehabilitation assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。