基于解剖结构的骨架动作识别模型,提升神经疾病诊断准确率
SkelMamba: A State Space Model for Efficient Skeleton Action Recognition of Neurological Disorders
- 按解剖结构分解骨骼运动为时空流,用通道分割高效捕捉动作特征
- 在多个公开数据集上最高提升3.2%准确率,计算量低于主流Transformer模型
- 专为神经疾病诊断设计,适合医疗影像与动作分析交叉研究者
我们提出一种基于状态空间模型(SSM)的新型骨架动作识别框架,采用解剖引导架构,在临床诊断与通用动作识别任务中均实现当前最优性能。该方法将骨骼运动分析分解为空间、时间及时空流,通过通道分割高效捕捉不同运动特征。在状态空间模型中引入结构化多方向扫描策略,有效捕捉局部关节交互与跨多个解剖部位的全局运动模式。这种解剖感知的分解方式显著增强对细微运动模式的识别能力,尤其适用于识别与神经系统疾病相关的步态异常。在公开基准数据集NTU RGB+D、NTU RGB+D 120和NW-UCLA上,本模型优于现有最先进方法,准确率最高提升3.2%,且计算复杂度低于先前领先的Transformer模型。此外,我们构建了一个新的医学数据集,用于基于运动的患者神经疾病分析,验证了该方法在自动化疾病诊断中的潜力。
原文摘要 · Abstract (English)
We introduce a novel state-space model (SSM)-based framework for skeleton-based human action recognition, with an anatomically-guided architecture that improves state-of-the-art performance in both clinical diagnostics and general action recognition tasks. Our approach decomposes skeletal motion analysis into spatial, temporal, and spatio-temporal streams, using channel partitioning to capture distinct movement characteristics efficiently. By implementing a structured, multi-directional scanning strategy within SSMs, our model captures local joint interactions and global motion patterns across multiple anatomical body parts. This anatomically-aware decomposition enhances the ability to identify subtle motion patterns critical in medical diagnosis, such as gait anomalies associated with neurological conditions. On public action recognition benchmarks, i.e., NTU RGB+D, NTU RGB+D 120, and NW-UCLA, our model outperforms current state-of-the-art methods, achieving accuracy improvements up to $3.2\%$ with lower computational complexity than previous leading transformer-based models. We also introduce a novel medical dataset for motion-based patient neurological disorder analysis to validate our method's potential in automated disease diagnosis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。