无需分段即可评估发音质量,提升语音诊断精度
Segmentation-free Goodness of Pronunciation
- 提出自对齐与全分段自由的发音评分方法
- 在多个数据集上实现领先水平的发音评估效果
- 适合语音学习系统与声学模型研究者使用
发音错误检测与诊断(MDD)是现代计算机辅助语言学习(CALL)系统的重要组成部分。现有基于音素级发音质量评分(GOP)的方法通常依赖于预先的语音分段,限制了准确性,并阻碍了使用基于CTC的声学模型进行评估。本文首先提出自对齐GOP(GOP-SA),使基于CTC训练的语音识别模型可用于MDD;随后定义更通用的无分段方法GOP-SF,综合考虑标准转录的所有可能分段。我们提供了GOP-SF的理论分析、数值稳定性处理方案及归一化方法,使其适用于不同时间峰值特性的声学模型。在CMU Kids和speechocean762数据集上进行了广泛实验,验证了方法的有效性,分析了其对声学模型峰值特性及目标音素上下文长度的依赖关系。与最新研究对比显示,基于本方法提取的特征在音素级发音评估中达到最优性能。
原文摘要 · Abstract (English)
Mispronunciation detection and diagnosis (MDD) is a significant part in modern computer-aided language learning (CALL) systems. Most systems implementing phoneme-level MDD through goodness of pronunciation (GOP), however, rely on pre-segmentation of speech into phonetic units. This limits the accuracy of these methods and the possibility to use modern CTC-based acoustic models for their evaluation. In this study, we first propose self-alignment GOP (GOP-SA) that enables the use of CTC-trained ASR models for MDD. Next, we define a more general segmentation-free method that takes all possible segmentations of the canonical transcription into account (GOP-SF). We give a theoretical account of our definition of GOP-SF, an implementation that solves potential numerical issues as well as a proper normalization which allows the use of acoustic models with different peakiness over time. We provide extensive experimental results on the CMU Kids and speechocean762 datasets comparing the different definitions of our methods, estimating the dependency of GOP-SF on the peakiness of the acoustic models and on the amount of context around the target phoneme. Finally, we compare our methods with recent studies over the speechocean762 data showing that the feature vectors derived from the proposed method achieve state-of-the-art results on phoneme-level pronunciation assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。