arXiv:2608.10206cs.AIcs.CL2026-08

通过年龄感知训练,轻量模型在儿童语音识别上超越大模型。

Edge Phoneme Recognition for Children's Speech through Age-Aware Training

  • 同时预测年龄和音素序列,提升儿童语音识别效果。
  • 9400万参数模型在特定数据集上超越3.17亿参数的WavLM Large。
  • 可部署于手机端,适合隐私敏感的儿童语音应用。

由于儿童语音训练数据稀缺及其独特特征,儿童语音中的音素检测一直存在困难。在一次音素检测竞赛中,我们发现训练一个轻量级模型同时预测学习者年龄和音素序列,使得一个9400万参数的模型在目标DrivenData数据分布上表现优于3.17亿参数的WavLM Large模型,并且与参数量达其90倍的竞赛集成模型相比,错误率(CER)仅相差约0.04。该方法催生了PhonemeTrainer应用,可在多数现代智能手机上运行,有望提升儿童语音的自动语音识别与发音辅助应用性能,同时带来边缘计算所具备的隐私与合规优势。

原文摘要 · Abstract (English)

Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.04 CER of competition ensembles with 90 times the parameters. This has enabled the creation of PhonemeTrainer, an application that can run on most modern cellular phones. This will ultimately enable better Automated Speech Recognition (ASR) and pronunciation helper apps for children's speech, with the privacy and compliance benefits that come with edge processing.

语音识别儿童语音边缘计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。