arXiv:2603.22371eess.IV2026-03

融合骨骼动态与临床步态特征,提升脑瘫严重程度评估准确率

Multimodal Fusion of Skeleton Dynamics and Clinical Gait Features for Video-Based Cerebral Palsy Severity Assessment

  • 双流架构分别建模骨骼动态和关键点导出的临床步态特征
  • 通过跨注意力融合,四分类准确率达70.86%,提升5.6个百分点
  • 结合医学可解释性与模型性能,适合医疗辅助诊断场景

基于视频的步态分析已成为评估脑瘫患儿运动功能障碍的有前景方法。然而,现有方法通常仅依赖姿态序列或手工提取的步态特征,难以同时捕捉时空运动模式与具有临床意义的生物力学信息。为此,我们提出一种多模态融合框架,将骨骼动态与贡献引导的临床步态特征相结合。首先,利用预训练的ST-GCN模型上的Grad-CAM分析,识别出最具判别性的身体关键点,为后续步态特征提取提供可解释基础。随后构建双流架构:一路使用ST-GCN建模骨骼动态,另一路编码基于关键点提取的步态特征。通过特征交叉注意力融合两路信息,实现四等级脑瘫运动严重程度分类准确率70.86%,相比基线提升5.6个百分点。结果表明,融合骨骼动态与临床步态描述符可同时提升预测性能与生物力学可解释性。

原文摘要 · Abstract (English)

Video-based gait analysis has become a promising approach for assessing motor impairment in children with cerebral palsy (CP). However, existing methods usually rely on either pose sequences or handcrafted gait features alone, making it difficult to simultaneously capture spatiotemporal motion patterns and clinically meaningful biomechanical information. To address this gap, we propose a multimodal fusion framework that integrates skeleton dynamics with contribution-guided clinically meaningful gait features. First, Grad-CAM analysis on a pre-trained ST-GCN backbone identified the most discriminative body keypoints, providing an interpretable basis for subsequent gait feature extraction. We then build a dual-stream architecture, with one stream modeling skeleton dynamics using ST-GCN and the other encoding gait geatures derived from the identified keypoints. By fusing the two streams through feature cross-attention improved four-level CP motor severity classification to 70.86%, outperforming the baseline by 5.6 percentage points. Overall, this work suggests that integrating skeleton dynamics with clinically meaningful gait descriptors can improve both prediction performance and biomechanical interpretability for video-based CP severity assessment.

脑瘫评估步态分析多模态融合可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。