arXiv:2504.14708cs.CVcs.AI2025-04被引 6

用细粒度特征提升肌电手势识别准确率

Time Frequency Analysis of EMG Signal for Gesture Recognition using Fine grained Features

  • 通过跨层互注意力融合浅层局部与深层语义特征
  • 在多个数据集上相比基线模型提升1.46%~9.36%
  • 适合肌电控制、假肢和人机交互场景

基于肌电(EMG)的手势识别将前臂肌肉活动转化为假肢控制、康复训练和人机交互的指令。本文提出一种新方法,利用细粒度分类并构建XMANet模型,通过浅层到深层CNN专家间的跨层互注意力,统一低层局部特征与高层语义信息。采用短时傅里叶变换(STFT)和小波变换(WT)生成堆叠频谱图与小波图作为输入,在Grabmyo和FORS EMG数据集上对比了ResNet50、DenseNet-121、MobileNetV3和EfficientNetB0。实验表明,使用STFT时,XMANet相较基线分别提升1.72%、4.38%、5.10%和2.53%;采用WT时,提升分别为1.57%、1.88%、1.46%和2.05%。在FORS数据集上,基于ResNet50的XMANet提升5.04%,基于DenseNet121和MobileNetV3的模型分别提升4.11%和2.81%。使用WT时,相对于各基线模型,性能提升达4.26%~9.36%。结果验证了XMANet在不同架构与信号处理方法下均具稳定增益,表明细粒度特征对精确、鲁棒的肌电分类具有显著潜力。

原文摘要 · Abstract (English)

Electromyography (EMG) based hand gesture recognition converts forearm muscle activity into control commands for prosthetics, rehabilitation, and human computer interaction. This paper proposes a novel approach to EMG-based hand gesture recognition that uses fine-grained classification and presents XMANet, which unifies low-level local and high level semantic cues through cross layer mutual attention among shallow to deep CNN experts. Using stacked spectrograms and scalograms derived from the Short Time Fourier Transform (STFT) and Wavelet Transform (WT), we benchmark XMANet against ResNet50, DenseNet-121, MobileNetV3, and EfficientNetB0. Experimental results on the Grabmyo dataset indicate that, using STFT, the proposed XMANet model outperforms the baseline ResNet50, EfficientNetB0, MobileNetV3, and DenseNet121 models with improvement of approximately 1.72%, 4.38%, 5.10%, and 2.53%, respectively. When employing the WT approach, improvements of around 1.57%, 1.88%, 1.46%, and 2.05% are observed over the same baselines. Similarly, on the FORS EMG dataset, the XMANet(ResNet50) model using STFT shows an improvement of about 5.04% over the baseline ResNet50. In comparison, the XMANet(DenseNet121) and XMANet(MobileNetV3) models yield enhancements of approximately 4.11% and 2.81%, respectively. Moreover, when using WT, the proposed XMANet achieves gains of around 4.26%, 9.36%, 5.72%, and 6.09% over the baseline ResNet50, DenseNet121, MobileNetV3, and EfficientNetB0 models, respectively. These results confirm that XMANet consistently improves performance across various architectures and signal processing techniques, demonstrating the strong potential of fine grained features for accurate and robust EMG classification.

肌电识别细粒度特征深度学习手势控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。