用数据自动发现发音错误模式,提升非母语语音识别准确率
Data-Driven Mispronunciation Pattern Discovery for Robust Speech Recognition
- 通过注意力图对齐非母语与母语音素,自动挖掘错误模式
- 对母语者数据提升5.7%,对韩语背景者提升12.8%
- 无需先验语言知识,适合实际场景的鲁棒语音识别
机器学习进步显著提升了语音识别性能,但非流利或带有口音的语音识别仍是难题。以往依赖规则的方法难以全面捕捉非母语发音错误。本文提出两种基于语音语料库的数据驱动方法,通过注意力图将非母语音素与母语音素对齐,实现了在母语英语数据集上5.7%的识别率提升,对非母语英语使用者(尤其是韩国说话者)达12.8%的提升。该方法为缺乏先验语言知识的场景提供了实用的鲁棒自动语音识别解决方案。
原文摘要 · Abstract (English)
Recent advancements in machine learning have significantly improved speech recognition, but recognizing speech from non-fluent or accented speakers remains a challenge. Previous efforts, relying on rule-based pronunciation patterns, have struggled to fully capture non-native errors. We propose two data-driven approaches using speech corpora to automatically detect mispronunciation patterns. By aligning non-native phones with their native counterparts using attention maps, we achieved a 5.7% improvement in speech recognition on native English datasets and a 12.8% improvement for non-native English speakers, particularly Korean speakers. Our method offers practical advancements for robust Automatic Speech Recognition (ASR) systems particularly for situations where prior linguistic knowledge is not applicable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。