用自监督学习和说话人分离,自动评估韩语幼儿发音准确度
Automated Pronunciation Evaluation for Korean Toddler Speech using Speech Diarization and Self-Supervised Learning
- 结合说话人分离与自监督语音表征,构建端到端评估流程
- 对53名2-5岁儿童的1938个音节标注,发音识别准确率达78.2%
- 适合语言发育障碍筛查、智能教育系统研发者参考
言语发音障碍影响约44%的韩国儿童语言障碍病例,但针对韩语幼儿语音的自动化评估工具仍不成熟。本文提出一种端到端的韩语幼儿发音评估流水线,融合神经说话人分离与自监督语音表示学习。构建了一个经伦理审批的全新语料库,包含53段2-5岁韩语儿童语音录音,由三位独立评审员标注,共获得1,190个辅音和748个元音的词级正确性标签。评估三种说话人分离模型,发现NeMo SortFormer在到达时间排序的Transformer架构下,实现88.69%的说话人数量准确率和33.04%的分离错误率(DER),有效处理了年轻女性照护者使用‘爱娇’语调与幼儿语音的声学混淆问题。在发音评分方面,比较三种自监督学习(SSL)主干网络,跨模型集成策略将辅音预测导向HuBERT-large,元音预测导向WavLM-large,分别取得0.720和0.845的准确率,平均达0.782。
原文摘要 · Abstract (English)
Speech sound disorders affect approximately 44% of Korean pediatric communication disorder cases, yet automated assessment tools for Korean toddler speech remain underdeveloped. This paper presents an end-to-end pipeline for automated pronunciation evaluation of Korean toddler speech, combining neural speaker diarization with self-supervised speech representation learning. We introduce a novel IRB-approved corpus of 53 recordings from Korean-speaking children aged 2-5 years. A subset of 53 subjects was annotated by three independent reviewers, yielding 1,190 consonant and 748 vowel word-level binary correctness labels. We evaluate three diarization models, finding that NeMo SortFormer achieves 88.69% speaker count accuracy and 33.04% diarization error rate (DER) owing to its arrival-time-sorted transformer architecture, which handles the acoustic confound between young female caregivers exhibiting aegyo and toddler speech. For pronunciation scoring, we compare three self-supervised learning (SSL) backbones across multiple pooling strategies. A cross-model ensemble routing consonant prediction to HuBERT-large and vowel prediction to WavLM-large achieves balanced accuracies of 0.720 and 0.845, with a mean of 0.782.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。