用最少传感器实现高精度手势识别,适合低成本假肢控制。
Time2Vec Transformer for Robust Gesture Recognition from Low-Density sEMG
- 引入可学习的时间嵌入,捕捉生物信号的随机时间扭曲。
- 在8名受试者上达95.7%的多主体分类准确率,优于传统模型。
- 仅需两次校准试验即可恢复性能,适合快速个性化适配。
精确且响应迅速的肌电假肢控制通常依赖复杂的密集多传感器阵列,限制了消费级应用。本文提出一种新型、数据高效的深度学习框架,仅用最少的传感器硬件即可实现精准控制。利用8名受试者的外部数据集,该方法采用针对稀疏双通道表面肌电(sEMG)优化的混合Transformer架构。不同于使用固定位置编码的标准结构,我们集成可学习的时间嵌入(Time2Vec),以捕捉生物信号固有的随机时间扭曲。此外,采用归一化加性融合策略,对齐空间与时间特征的潜在分布,避免标准实现中的破坏性干扰。通过两阶段课程学习协议,确保在数据稀缺情况下的鲁棒特征提取。所提架构在10类动作集上达到95.7% ± 0.20%的多主体F1分数,显著优于标准Transformer和循环CNN-LSTM模型。架构优化表明,空间与时间维度模型容量的均衡分配带来最高稳定性。尽管直接迁移至未见受试者导致准确率下降至21.0% ± 2.98%,但仅需每动作两次校准试验,性能即可恢复至96.9% ± 0.52%。本工作验证了高保真时间嵌入可补偿低空间分辨率,挑战了高密度传感的必要性。该框架为下一代可快速个性化、成本低廉的假肢接口提供了稳健蓝图。
原文摘要 · Abstract (English)
Accurate and responsive myoelectric prosthesis control typically relies on complex, dense multi-sensor arrays, which limits consumer accessibility. This paper presents a novel, data-efficient deep learning framework designed to achieve precise and accurate control using minimal sensor hardware. Leveraging an external dataset of 8 subjects, our approach implements a hybrid Transformer optimized for sparse, two-channel surface electromyography (sEMG). Unlike standard architectures that use fixed positional encodings, we integrate Time2Vec learnable temporal embeddings to capture the stochastic temporal warping inherent in biological signals. Furthermore, we employ a normalized additive fusion strategy that aligns the latent distributions of spatial and temporal features, preventing the destructive interference common in standard implementations. A two-stage curriculum learning protocol is utilized to ensure robust feature extraction despite data scarcity. The proposed architecture achieves a state-of-the-art multi-subject F1-score of 95.7% $\pm$ 0.20% for a 10-class movement set, statistically outperforming both a standard Transformer with fixed encodings and a recurrent CNN-LSTM model. Architectural optimization reveals that a balanced allocation of model capacity between spatial and temporal dimensions yields the highest stability. Furthermore, while direct transfer to a new unseen subject led to poor accuracy due to domain shifts, a rapid calibration protocol utilizing only two trials per gesture recovered performance from 21.0% $\pm$ 2.98% to 96.9% $\pm$ 0.52%. By validating that high-fidelity temporal embeddings can compensate for low spatial resolution, this work challenges the necessity of high-density sensing. The proposed framework offers a robust, cost-effective blueprint for next-generation prosthetic interfaces capable of rapid personalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。