用深度学习实现实时手语识别,准确率达88.23%。
Continuous Sign Language Recognition System using Deep Learning with MediaPipe Holistic
- 基于MediaPipe Holistic提取面部、手部和身体关键点,构建数据集。
- 采用LSTM模型实现连续手语识别,实时识别准确率88.23%。
- 适用于残障人士沟通辅助,尤其适合印度手语场景。
手语是听障人士通过手势、面部表情和身体动作进行交流的视觉语言。全球约有300种手语,如美国手语(ASL)、中国手语(CSL)、印度手语(ISL)等,其表达方式与当地口语相关。手语不依赖助动词,仅靠视觉符号传递信息。由于掌握手语的人群有限,听障人士难以与他人顺畅交流。本研究提出一种基于深度学习的连续手语识别(SLR)系统,利用MediaPipe Holistic管道捕捉人脸、手部和身体关键点,构建印度手语(ISL)基础数据集,并训练长短期记忆网络(LSTM)模型。系统可实现实时手语识别,准确率达到88.23%。
原文摘要 · Abstract (English)
Sign languages are the language of hearing-impaired people who use visuals like the hand, facial, and body movements for communication. There are different signs and gestures representing alphabets, words, and phrases. Nowadays approximately 300 sign languages are being practiced worldwide such as American Sign Language (ASL), Chinese Sign Language (CSL), Indian Sign Language (ISL), and many more. Sign languages are dependent on the vocal language of a place. Unlike vocal or spoken languages, there are no helping words in sign language like is, am, are, was, were, will, be, etc. As only a limited population is well-versed in sign language, this lack of familiarity of sign language hinders hearing-impaired people from communicating freely and easily with everyone. This issue can be addressed by a sign language recognition (SLR) system which has the capability to translate the sign language into vocal language. In this paper, a continuous SLR system is proposed using a deep learning model employing Long Short-Term Memory (LSTM), trained and tested on an ISL primary dataset. This dataset is created using MediaPipe Holistic pipeline for tracking face, hand, and body movements and collecting landmarks. The system recognizes the signs and gestures in real-time with 88.23% accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。