无需音素对齐,用弱监督模型实现跨语言语音质量评估
Goodness-of-pronunciation without phoneme time alignment
- 通过混淆网络映射声学假设获取音素后验
- 采用词级语速与帧级特征融合,性能接近标准方法
- 适合低资源语言语音评估,无需音素时间对齐
在语音评估中,自动语音识别(ASR)模型通常需计算时间边界和音素后验。但受限于训练数据,当前ASR难以扩展至低资源语言。开源弱监督模型虽可覆盖多种语言,但其帧异步且非音素级,阻碍特征提取。本文提出新方法:通过将ASR预测映射到音素混淆网络生成音素后验,使用词级而非音素级的说话速率和时长,并结合交叉注意力架构融合音素与帧级特征,避免音素时间对齐。该方法在英语SpeechOcean762和低资源泰米尔语数据集上表现与标准帧同步特征相当。
原文摘要 · Abstract (English)
In speech evaluation, an Automatic Speech Recognition (ASR) model often computes time boundaries and phoneme posteriors for input features. However, limited data for ASR training hinders expansion of speech evaluation to low-resource languages. Open-source weakly-supervised models are capable of ASR over many languages, but they are frame-asynchronous and not phonemic, hindering feature extraction for speech evaluation. This paper proposes to overcome incompatibilities for feature extraction with weakly-supervised models, easing expansion of speech evaluation to low-resource languages. Phoneme posteriors are computed by mapping ASR hypotheses to a phoneme confusion network. Word instead of phoneme-level speaking rate and duration are used. Phoneme and frame-level features are combined using a cross-attention architecture, obviating phoneme time alignment. This performs comparably with standard frame-synchronous features on English speechocean762 and low-resource Tamil datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。