用摩擦纳米发电机手套+频域特征,提升手语识别准确率
Development of ML model for triboelectric nanogenerator based sign language detection system

- 通过多传感器频域特征融合,构建并行卷积-循环网络
- 达93.33%准确率,较最优传统模型提升23个百分点
- 适合可穿戴设备手语识别,尤其关注抗干扰与低功耗
手语识别对弥合聋哑人与听力人群沟通鸿沟至关重要。基于视觉的方法存在遮挡、计算成本高和物理限制问题。本文对比了机器学习与深度学习模型在自研摩擦纳米发电机(TENG)传感器手套上的表现。利用五个弯曲传感器的多变量时序数据,在11个手势类别(数字1-5,字母A-F)上评估了传统机器学习算法、前馈神经网络、LSTM时间模型及多传感器MFCC-CNN-LSTM架构。提出的MFCC-CNN-LSTM架构将各传感器的频域特征通过独立卷积分支处理后融合,实现93.33%准确率与95.56%精确率,较最佳传统模型(随机森林:70.38%)提升23个百分点。消融实验表明,50步长窗口在时序上下文与训练数据量间取得平衡,准确率达84.13%,而100步长仅58.06%。MFCC特征提取将时序变化映射为执行速度无关的谱表示,时间扭曲与噪声注入等数据增强方法对泛化能力至关重要。结果表明,频域特征与并行多传感器处理架构显著优于经典算法与时域深度学习方法,有助于可穿戴辅助技术发展。
原文摘要 · Abstract (English)
Sign language recognition (SLR) is vital for bridging communication gaps between deaf and hearing communities. Vision-based approaches suffer from occlusion, computational costs, and physical constraints. This work presents a comparison of machine learning (ML) and deep learning models for a custom triboelectric nanogenerator (TENG)-based sensor glove. Utilizing multivariate time-series data from five flex sensors, the study benchmarks traditional ML algorithms, feedforward neural networks, LSTM-based temporal models, and a multi-sensor MFCC CNN-LSTM architecture across 11 sign classes (digits 1-5, letters A-F). The proposed MFCC CNN-LSTM architecture processes frequency-domain features from each sensor through independent convolutional branches before fusion. It achieves 93.33% accuracy and 95.56% precision, a 23-point improvement over the best ML algorithm (Random Forest: 70.38%). Ablation studies reveal 50-timestep windows offer a tradeoff between temporal context and training data volume, yielding 84.13% accuracy compared to 58.06% with 100-timestep windows. MFCC feature extraction maps temporal variations to execution-speed-invariant spectral representations, and data augmentation methods (time warping, noise injection) are essential for generalization. Results demonstrate that frequency-domain feature representations combined with parallel multi-sensor processing architectures offer enhancement over classical algorithms and time-domain deep learning for wearable sensor-based gesture recognition. This aids assistive technology development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。