首个针对德语的唇读深度学习模型,准确率达87%。
Development and evaluation of a deep learning algorithm for German word recognition from lip movements
- 用3D卷积网络与门控循环单元结合,从唇部动作识别德语词汇。
- 对已知说话人识别准确率最高达87%,未知说话人仍达63%。
- 首次实现德语唇读高精度,可扩展至更多词汇类别。
阅读唇语时,许多用户依赖说话者嘴唇动作的视觉信息,但该方式极易出错。基于人工神经网络的人工智能唇读算法显著提升词汇识别准确率,但此前缺乏针对德语的解决方案。研究选取1806段仅含一名德语说话者的视频,将其分割为词段,并通过语音识别软件标注词类。在包含32名说话人、38,391个视频片段的语料中,使用18个音节分明、视觉可区分的德语词汇训练并验证神经网络。对比了3D卷积神经网络、门控循环单元(GRU)及两者结合的GRUConv模型,以及不同图像区域和色彩空间的效果。经过5000次训练迭代评估准确率。结果显示,色彩空间差异对分类准确率影响不显著(69%~72%)。仅截取唇部区域时,准确率达70%,显著高于包含整张脸的34%。采用GRUConv模型,在已知说话人情况下最高准确率达87%,未知说话人验证中达63%。该首次为德语设计的唇读神经网络展现出极高准确率,与英语系统相当,且具备泛化潜力,未来可拓展至更多词类。
原文摘要 · Abstract (English)
When reading lips, many people benefit from additional visual information from the lip movements of the speaker, which is, however, very error prone. Algorithms for lip reading with artificial intelligence based on artificial neural networks significantly improve word recognition but are not available for the German language. A total of 1806 video clips with only one German-speaking person each were selected, split into word segments, and assigned to word classes using speech-recognition software. In 38,391 video segments with 32 speakers, 18 polysyllabic, visually distinguishable words were used to train and validate a neural network. The 3D Convolutional Neural Network and Gated Recurrent Units models and a combination of both models (GRUConv) were compared, as were different image sections and color spaces of the videos. The accuracy was determined in 5000 training epochs. Comparison of the color spaces did not reveal any relevant different correct classification rates in the range from 69% to 72%. With a cut to the lips, a significantly higher accuracy of 70% was achieved than when cut to the entire speaker's face (34%). With the GRUConv model, the maximum accuracies were 87% with known speakers and 63% in the validation with unknown speakers. The neural network for lip reading, which was first developed for the German language, shows a very high level of accuracy, comparable to English-language algorithms. It works with unknown speakers as well and can be generalized with more word classes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。