用AI自动评估低资源语言儿童口语流利度,准确率超现有方法
Automated evaluation of children's speech fluency for low-resource languages
- 融合微调多语言语音识别与GPT分类器,结合语音错误率等客观指标
- 在泰米尔语和马来语儿童语音数据上准确率显著优于随机森林等方法
- 适合教育科技、低资源语言语音评估场景,尤其关注儿童语言发展
在主流语言中,儿童口语流利度的评估已较为成熟,但在低资源语言中仍具挑战。本文提出一种自动评估系统,结合微调的多语言语音识别模型、客观指标提取阶段以及生成式预训练变换器(GPT)网络。客观指标包括音素错误率、单词错误率、语速及语音-停顿时长比。这些指标由基于少量人工标注样本引导的GPT分类器进行解释,以评分流利度。我们在泰米尔语和马来语儿童语音数据集上评估该系统,并与随机森林、XGBoost及直接使用ChatGPT-4o从语音输入预测流利度的方法进行对比。结果表明,所提方法在分类性能上显著优于多模态GPT及其他方法。
原文摘要 · Abstract (English)
Assessment of children's speaking fluency in education is well researched for majority languages, but remains highly challenging for low resource languages. This paper proposes a system to automatically assess fluency by combining a fine-tuned multilingual ASR model, an objective metrics extraction stage, and a generative pre-trained transformer (GPT) network. The objective metrics include phonetic and word error rates, speech rate, and speech-pause duration ratio. These are interpreted by a GPT-based classifier guided by a small set of human-evaluated ground truth examples, to score fluency. We evaluate the proposed system on a dataset of children's speech in two low-resource languages, Tamil and Malay and compare the classification performance against Random Forest and XGBoost, as well as using ChatGPT-4o to predict fluency directly from speech input. Results demonstrate that the proposed approach achieves significantly higher accuracy than multimodal GPT or other methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。