用深度学习实现手语实时翻译与打分,提升学习效率
Learning Sign Language Representation using CNN LSTM, 3DCNN, CNN RNN LSTM and CCN TD
- 对比CNN、3DCNN等模型,3DCNN擅长捕捉时空特征
- 在特立尼达和多巴哥手语数据集上达到91%准确率
- 适合手语教学系统开发与智能评估工具研究者
现有手语学习应用多依赖示范,需教师手动回看视频确认动作正确性。本文探索实现实时视频手语翻译与初学者手语准确度评分的算法,要求模型能有效识别和处理空间与时间特征。通过在特立尼达和多巴哥手语(TTSL)及美国手语(ASL)两个数据集上比较CNN、3DCNN、CNN-LSTM、CNN-RNN-LSTM等主流算法,发现3DCNN表现最佳,在TTSL数据集上准确率达91%,在ASL数据集上为83%。研究旨在为构建智能化手语教学系统提供最优神经网络方案。
原文摘要 · Abstract (English)
Existing Sign Language Learning applications focus on the demonstration of the sign in the hope that the student will copy a sign correctly. In these cases, only a teacher can confirm that the sign was completed correctly, by reviewing a video captured manually. Sign Language Translation is a widely explored field in visual recognition. This paper seeks to explore the algorithms that will allow for real-time, video sign translation, and grading of sign language accuracy for new sign language users. This required algorithms capable of recognizing and processing spatial and temporal features. The aim of this paper is to evaluate and identify the best neural network algorithm that can facilitate a sign language tuition system of this nature. Modern popular algorithms including CNN and 3DCNN are compared on a dataset not yet explored, Trinidad and Tobago Sign Language as well as an American Sign Language dataset. The 3DCNN algorithm was found to be the best performing neural network algorithm from these systems with 91% accuracy in the TTSL dataset and 83% accuracy in the ASL dataset.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。