对比三种端到端模型,提升儿童挪威语发音评估准确率
Comparison of End-to-end Speech Assessment Models for the NOCASA 2025 Challenge
- 用语音对齐的CTC计算发音质量特征,融合到新模型中
- 最佳模型在未加权召回率和平均绝对误差上超越基线
- 适合关注语音评估与教育技术的研究者
本文分析了为NOCASA 2025挑战赛开发的三种端到端模型,旨在对学习挪威语的儿童进行词级发音自动评估。模型包括编码器-解码器孪生结构(E2E-R)、基于预训练wav2vec2.0表示的前缀微调直接分类模型,以及一种结合无对齐发音质量(GOP)特征(通过CTC计算)的新模型。我们引入了一种针对未加权平均召回率和平均绝对误差优化的加权序数交叉熵损失。在所考察的方法中,基于GOP-CTC的模型表现最优,显著优于挑战赛基线,并取得了最高排行榜分数。
原文摘要 · Abstract (English)
This paper presents an analysis of three end-to-end models developed for the NOCASA 2025 Challenge, aimed at automatic word-level pronunciation assessment for children learning Norwegian as a second language. Our models include an encoder-decoder Siamese architecture (E2E-R), a prefix-tuned direct classification model leveraging pretrained wav2vec2.0 representations, and a novel model integrating alignment-free goodness-of-pronunciation (GOP) features computed via CTC. We introduce a weighted ordinal cross-entropy loss tailored for optimizing metrics such as unweighted average recall and mean absolute error. Among the explored methods, our GOP-CTC-based model achieved the highest performance, substantially surpassing challenge baselines and attaining top leaderboard scores.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。