用轻量级混合储池计算模型实现低成本手语视频识别
A Low-Cost Hybrid Reservoir Computing Model for Isolated Sign Language Video Recognition

- 结合人体关键点与混合储池架构提取动态特征
- 在WLASL100数据集上达61.12%准确率,训练仅需数秒
- 适合边缘设备部署,比深度学习更省算力
手语识别(SLR)有助于听障人士与健听者沟通。尽管深度学习(DL)在SLR中表现优异,但其高计算成本限制了在边缘设备上的部署。为此,我们提出一种基于轻量级储池计算(RC)的手语识别方法。该方法利用MediaPipe提取身体和手部关键点,捕捉手势的空间与时间动态;随后通过融合深度储池计算(DRC)与双向储池计算(BRC)的混合储池计算(HRC)架构,将输入转换为高维动态表示;最终由岭回归模型映射至类别标签。该方法在词级美国手语100(WLASL100)视频数据集上达到Top-1、Top-5、Top-10准确率分别为61.12%、86.05%、92.56%,性能媲美深度学习方法。此外,由于储池计算的轻量化特性,训练时间仅需数秒,远低于如Bi-GRU等深度学习方法。
原文摘要 · Abstract (English)
Sign language recognition (SLR) enhances communication between hearing and hearing-impaired individuals. Although deep learning (DL) has achieved promising performance in SLR, its high computational cost limits deployment on edge devices. To address this challenge, we propose a lightweight reservoir computing (RC)-based approach for SLR. In the proposed method, MediaPipe extracts body and hand keypoints to capture the spatial and temporal dynamics of gestures. These keypoints are then processed by a hybrid reservoir computing (HRC) architecture that combines deep reservoir computing (DRC) and bidirectional reservoir computing (BRC), transforming the input into a high-dimensional dynamic representation. A ridge regression model maps the final HRC state to class labels. This HRC-based SLR method achieved Top-1, Top-5, and Top-10 accuracies of 61.12%, 86.05%, and 92.56%, respectively, on the Word-Level American Sign Language 100 (WLASL100) video dataset, demonstrating competitive performance compared to deep learning-based approaches. Additionally, due to the lightweight nature of RC, the training time was drastically reduced to only a few seconds compared with DL-based methods such as Bi-GRU.This method offers low computational cost, showing its potential for deployment on edge devices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。