arXiv:2510.22011cs.CVcs.AI2025-10

用混合模型实现手语实时识别,准确率达92%

Reconnaissance Automatique des Langues des Signes : Une Approche Hybridée CNN-LSTM Basée sur Mediapipe

  • 结合CNN与LSTM,通过MediaPipe提取手势关键点
  • 平均准确率92%,对'你好''谢谢'等动作识别效果好
  • 适合残障辅助、教育医疗等场景应用

手语在聋人社群沟通中至关重要,但常被忽视,限制了其获取医疗、教育等基本服务的机会。本研究提出一种基于混合CNN-LSTM架构的手语自动识别系统,利用MediaPipe进行手势关键点提取,采用Python、TensorFlow和Streamlit开发,支持实时手势翻译。结果显示平均准确率达92%,对'Hello'和'Thank you'等明确手势表现优异;但部分视觉相似手势如'Call'与'Yes'仍存在混淆。该工作为医疗、教育及公共服务等领域提供了有前景的应用方向。

原文摘要 · Abstract (English)

Sign languages play a crucial role in the communication of deaf communities, but they are often marginalized, limiting access to essential services such as healthcare and education. This study proposes an automatic sign language recognition system based on a hybrid CNN-LSTM architecture, using Mediapipe for gesture keypoint extraction. Developed with Python, TensorFlow and Streamlit, the system provides real-time gesture translation. The results show an average accuracy of 92\%, with very good performance for distinct gestures such as ``Hello'' and ``Thank you''. However, some confusions remain for visually similar gestures, such as ``Call'' and ``Yes''. This work opens up interesting perspectives for applications in various fields such as healthcare, education and public services.

手语识别CNN-LSTMMediaPipe实时翻译

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。