首个端到端西班牙语连续唇读系统,性能显著超越现有方法。
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
- 采用混合CTC/注意力架构,实现端到端唇读建模。
- 在两个不同数据集上达到当前最优结果,指标显著提升。
- 提供完整分析与基准,适合多语言视觉语音研究者参考。
视觉语音识别仍是开放性研究问题,需应对视觉模糊、说话人差异及静音建模等挑战。得益于大规模数据库和强大注意力机制,该领域近年取得显著进展。目前,除英语外的多种语言也成为研究热点。本文在西班牙语连续唇读方面取得重要突破:提出基于混合CTC/注意力架构的端到端系统,在两个性质不同的语料库上开展实验,均达到当前最优性能,显著优于此前最佳结果。此外,进行了详尽的消融研究,分析各组件对识别效果的影响;通过严谨的错误分析,探究影响系统学习的关键因素。最后,构建了一个新的西班牙语唇读基准。代码与训练模型已公开于 https://github.com/david-gimeno/evaluating-end2end-spanish-lipreading。
原文摘要 · Abstract (English)
Visual speech recognition remains an open research problem where different challenges must be considered by dispensing with the auditory sense, such as visual ambiguities, the inter-personal variability among speakers, and the complex modeling of silence. Nonetheless, recent remarkable results have been achieved in the field thanks to the availability of large-scale databases and the use of powerful attention mechanisms. Besides, multiple languages apart from English are nowadays a focus of interest. This paper presents noticeable advances in automatic continuous lipreading for Spanish. First, an end-to-end system based on the hybrid CTC/Attention architecture is presented. Experiments are conducted on two corpora of disparate nature, reaching state-of-the-art results that significantly improve the best performance obtained to date for both databases. In addition, a thorough ablation study is carried out, where it is studied how the different components that form the architecture influence the quality of speech recognition. Then, a rigorous error analysis is carried out to investigate the different factors that could affect the learning of the automatic system. Finally, a new Spanish lipreading benchmark is consolidated. Code and trained models are available at https://github.com/david-gimeno/evaluating-end2end-spanish-lipreading.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。