构建首个科萨语儿童阅读评估数据集,助力低资源语言自动识读
An End-to-End Approach for Child Reading Assessment in the Xhosa Language
- 基于EGRA体系构建科萨语儿童语音数据集,支持多标注与独立验证
- wav2vec 2.0在多类别联合训练下性能显著提升,尤其在小样本场景
- 为低资源语言儿童阅读评估提供可复现的端到端解决方案,适合教育AI研究者
儿童读写能力是影响其未来生活轨迹的重要因素,尤其在中低收入群体中更需针对性干预。阅读评估有助于衡量教育项目成效,而人工智能可作为经济高效的辅助工具。然而,针对低资源语言的儿童语音自动读写评估面临数据稀缺与儿童嗓音特性复杂等挑战。本研究聚焦南非使用的科萨语,构建首个包含十项词汇与字母的儿童语音数据集,该数据集源自早期阅读评估(EGRA)体系,每段录音经多位标注者在线标注,并由独立评审员对子样本进行验证。我们使用三个先进的端到端模型(wav2vec 2.0、HuBERT、Whisper)对该数据集进行评估。结果表明,模型性能受训练数据量与类别平衡性显著影响,这对低成本大规模数据收集具有关键意义。此外,实验显示wav2vec 2.0在多类别联合训练下表现更优,即使样本有限。
原文摘要 · Abstract (English)
Child literacy is a strong predictor of life outcomes at the subsequent stages of an individual's life. This points to a need for targeted interventions in vulnerable low and middle income populations to help bridge the gap between literacy levels in these regions and high income ones. In this effort, reading assessments provide an important tool to measure the effectiveness of these programs and AI can be a reliable and economical tool to support educators with this task. Developing accurate automatic reading assessment systems for child speech in low-resource languages poses significant challenges due to limited data and the unique acoustic properties of children's voices. This study focuses on Xhosa, a language spoken in South Africa, to advance child speech recognition capabilities. We present a novel dataset composed of child speech samples in Xhosa. The dataset is available upon request and contains ten words and letters, which are part of the Early Grade Reading Assessment (EGRA) system. Each recording is labeled with an online and cost-effective approach by multiple markers and a subsample is validated by an independent EGRA reviewer. This dataset is evaluated with three fine-tuned state-of-the-art end-to-end models: wav2vec 2.0, HuBERT, and Whisper. The results indicate that the performance of these models can be significantly influenced by the amount and balancing of the available training data, which is fundamental for cost-effective large dataset collection. Furthermore, our experiments indicate that the wav2vec 2.0 performance is improved by training on multiple classes at a time, even when the number of available samples is constrained.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。