挑战用AI评估非母语儿童单字发音,助力语言学习游戏
Non-native Children's Automatic Speech Assessment Challenge (NOCASA)
- 构建儿童二语发音评估数据集,含44名孩子10334条录音
- 基于205个挪威词汇的评分数据,标注等级1至5星
- 提供基准模型,最高识别准确率达36.37%(UAR)
本文介绍IEEE MLSP 2025会议中的‘非母语儿童自动发音评估’(NOCASA)数据竞赛。该挑战要求参赛者开发系统,用于评估年幼儿童第二语言(L2)学习者在游戏化发音训练应用中对单字发音的准确性。主要难点包括训练数据有限及发音水平类别严重不均衡。为加速开发,我们发布了伪匿名训练数据集TeflonNorL2,包含44名说话者对205个挪威语单词的10,334条录音,每条由人工按1至5星评分。此外,还发布了两个预训练基线系统:基于ComParE_16声学特征的SVM分类器和多任务wav2vec 2.0模型。后者在测试集上表现最佳,未加权平均召回率(UAR)达36.37%。
原文摘要 · Abstract (English)
This paper presents the "Non-native Children's Automatic Speech Assessment" (NOCASA) - a data competition part of the IEEE MLSP 2025 conference. NOCASA challenges participants to develop new systems that can assess single-word pronunciations of young second language (L2) learners as part of a gamified pronunciation training app. To achieve this, several issues must be addressed, most notably the limited nature of available training data and the highly unbalanced distribution among the pronunciation level categories. To expedite the development, we provide a pseudo-anonymized training data (TeflonNorL2), containing 10,334 recordings from 44 speakers attempting to pronounce 205 distinct Norwegian words, human-rated on a 1 to 5 scale (number of stars that should be given in the game). In addition to the data, two already trained systems are released as official baselines: an SVM classifier trained on the ComParE_16 acoustic feature set and a multi-task wav2vec 2.0 model. The latter achieves the best performance on the challenge test set, with an unweighted average recall (UAR) of 36.37%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。