用YOLOv8与CNN实现实时手语识别与文本转手语,打通聋哑人沟通障碍
Enhancing Bidirectional Sign Language Communication: Integrating YOLOv8 and NLP for Real-Time Gesture Recognition & Translation
- 结合YOLOv8与CNN实时提取视频中手语的时空特征
- 实现从文本到手语动作的实时生成与展示
- 首个支持双向实时手语交互的ASL系统
本研究旨在通过实时摄像头画面处理美国手语(ASL)数据,并将其转换为文本信息。同时,构建一个可将文本实时转化为手语动作的框架,以帮助听力障碍者跨越语言障碍。研究采用YOLO模型和卷积神经网络(CNN)进行手语识别,其中YOLO模型在无需先验知识的情况下实时提取原始视频流中的判别性时空特征,避免设计缺陷;CNN模型也同步运行于实时场景下完成手语检测。此外,提出一种新方法:输入句子后,系统识别关键词并实时生成对应手语动作视频。据我们所知,这是首个在真实场景下实现双向实时美国手语交互的研究。
原文摘要 · Abstract (English)
The primary concern of this research is to take American Sign Language (ASL) data through real time camera footage and be able to convert the data and information into text. Adding to that, we are also putting focus on creating a framework that can also convert text into sign language in real time which can help us break the language barrier for the people who are in need. In this work, for recognising American Sign Language (ASL), we have used the You Only Look Once(YOLO) model and Convolutional Neural Network (CNN) model. YOLO model is run in real time and automatically extracts discriminative spatial-temporal characteristics from the raw video stream without the need for any prior knowledge, eliminating design flaws. The CNN model here is also run in real time for sign language detection. We have introduced a novel method for converting text based input to sign language by making a framework that will take a sentence as input, identify keywords from that sentence and then show a video where sign language is performed with respect to the sentence given as input in real time. To the best of our knowledge, this is a rare study to demonstrate bidirectional sign language communication in real time in the American Sign Language (ASL).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。