用六轴加速度计实现高精度无声语音识别,准确率达97.17%。
Silent Speech Sentence Recognition with Six-Axis Accelerometers using Conformer and CTC Algorithm
- 结合Conformer与CTC算法,从面部运动信号中解码无声语句
- 在数据库词表约束下,句子识别准确率达97.17%
- 适合语音障碍者使用,为无声交互提供新硬件方案
无声语音接口(SSI)正被积极研发,以帮助因沟通障碍而长期受困、生活质量下降的群体。然而,由于省略和连读现象,无声语句难以分割与识别。本文提出一种新型无声语句识别方法,利用六轴加速度计采集的面部运动信号,转换为转录的词与句子。采用基于Conformer的神经网络与连接时序分类(CTC)算法,实现上下文理解,并将非声学信号转化为词序列,仅依赖数据库中的候选词。测试结果显示,该方法在句子识别上达到97.17%的准确率,显著超越现有方法(典型准确率为85%-95%),证明加速度计作为高精度无声语音识别的可行模态。
原文摘要 · Abstract (English)
Silent speech interfaces (SSI) are being actively developed to assist individuals with communication impairments who have long suffered from daily hardships and a reduced quality of life. However, silent sentences are difficult to segment and recognize due to elision and linking. A novel silent speech sentence recognition method is proposed to convert the facial motion signals collected by six-axis accelerometers into transcribed words and sentences. A Conformer-based neural network with the Connectionist-Temporal-Classification algorithm is used to gain contextual understanding and translate the non-acoustic signals into words sequences, solely requesting the constituent words in the database. Test results show that the proposed method achieves a 97.17% accuracy in sentence recognition, surpassing the existing silent speech recognition methods with a typical accuracy of 85%-95%, and demonstrating the potential of accelerometers as an available SSI modality for high-accuracy silent speech sentence recognition.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。