用深度学习将手语实时转为语音,助听障人士顺畅交流。
Real-Time Sign Language Gestures to Speech Transcription using Deep Learning
- 用CNN模型识别摄像头捕捉的手势
- 在Sign Language MNIST数据集上准确率高
- 适合需要实时手语翻译的无障碍场景
沟通障碍对听障和言语障碍者构成重大挑战,常限制其在日常环境中的有效互动。本项目提出一种基于深度学习的实时辅助技术,将手语手势转换为文本与语音输出。系统采用卷积神经网络(CNN)在Sign Language MNIST数据集上训练,可实时捕获并准确分类摄像头采集的手势。检测到的姿势即时转化为对应语义,并通过文本转语音合成生成可听语音,实现无缝沟通。大量实验表明模型具备高准确率与稳健的实时性能,虽有轻微延迟,但整体表现可靠,具有良好的实用性,可作为提升听障用户社会融入度与自主性的便捷工具。
原文摘要 · Abstract (English)
Communication barriers pose significant challenges for individuals with hearing and speech impairments, often limiting their ability to effectively interact in everyday environments. This project introduces a real-time assistive technology solution that leverages advanced deep learning techniques to translate sign language gestures into textual and audible speech. By employing convolution neural networks (CNN) trained on the Sign Language MNIST dataset, the system accurately classifies hand gestures captured live via webcam. Detected gestures are instantaneously translated into their corresponding meanings and transcribed into spoken language using text-to-speech synthesis, thus facilitating seamless communication. Comprehensive experiments demonstrate high model accuracy and robust real-time performance with some latency, highlighting the system's practical applicability as an accessible, reliable, and user-friendly tool for enhancing the autonomy and integration of sign language users in diverse social settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。