arXiv:2606.31508cs.CLcs.SD2026-06

为巴马拉语儿童阅读构建开源语音识别系统,助力非洲语言识字评估

Building an ASR Solution for Training and Assessing Children's Reading

  • 基于60名儿童55小时语音数据,构建巴马拉语儿童阅读公开基准
  • 优化模型将错误率从42%降至22%,显著优于现有架构
  • 适用于教育机构、语言技术研究者及非洲语言数字化项目

针对大多数非洲语言(包括巴马拉语)的儿童阅读自动语音识别仍不成熟,尽管其在可重复识字评估中具有重要价值。本文提出一个开源系统,涵盖实地数据采集、基准构建、模型适配、阅读应用开发与课堂验证全流程。通过移动端应用收集了60名儿童共计55小时原始朗读语音,建立了公开的巴马拉语儿童阅读评估基准。对比实验采用适应巴马拉语的Fast-Conformer框架Soloni(含TDT与CTC解码器)与紧凑型卷积架构QuartzNet。最佳Soloni模型将词错误率(WER)从0.42降至0.22,字符错误率(CER)从0.15降至0.08,在孤立测试中显著优于QuartzNet。实验还发现:重复朗读对QuartzNet提升明显,对Soloni仅带来边际收益;SpecAugment训练增强未超越最优无增强配置。分项分析指出,10岁以下儿童是残余错误主要来源,提示需针对性采集更小年龄组数据。十次课堂试用表明该应用具备持续使用潜力。

原文摘要 · Abstract (English)

Automatic speech recognition for children's reading remains underdeveloped for most African languages, including Bambara, despite its potential value for reproducible literacy assessment. We present an open-source system for assessing children's reading in Bambara, developed through an end-to-end process linking field data collection, benchmark construction, model adaptation, a reading application, and classroom validation. A mobile collection and assessment app was used to collect 55 hours of raw reading speech from 60 children, from which we construct a public benchmark for Bambara child-reading assessment. Fine-tuning experiments compare Soloni, a Bambara-adapted Fast-Conformer ASR framework with TDT and CTC decoders, with QuartzNet, a compact convolutional ASR architecture. The best Soloni model reduces WER from 0.42 to 0.22 and CER from 0.15 to 0.08, substantially outperforming QuartzNet on the isolated benchmark. The experiments further show that repeated readings of the same texts provide architecture-dependent benefits: they substantially improve QuartzNet but add only marginal gains for Soloni, while SpecAugment regulates training without exceeding the best unaugmented configuration. Disaggregated analysis identifies children under 10 as the main source of residual errors, motivating targeted collection from younger readers. Ten classroom trials supported continued use of the application.

语音识别儿童阅读非洲语言教育科技

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。