对比不同训练策略,发现儿童语音识别需专用数据与模型。
Benchmarking Training Paradigms, Dataset Composition, and Model Scaling for Child ASR in ESPnet
- 用儿童语音数据从头训练,可缓解预训练模型对成人的偏见。
- 模型参数量达10亿时性能最佳,继续扩大无明显提升。
- 适合关注儿童语音识别的开发者与研究者参考。
尽管语音识别技术不断进步,儿童语音识别仍因声学差异大、标注数据少而困难。虽常将成人语音模型微调用于儿童语音,但与从头训练的对比研究仍不足。本文在ESPnet框架下,比较了多种数据集、自监督学习表示(WavLM、XEUS)及解码器架构的从头训练效果。结果表明,自监督表示存在成人语音偏向,使用儿童语音数据从头训练可有效缓解此偏差。同时,模型规模分析显示,性能随参数量增加持续提升至10亿参数后趋于饱和。年龄相关语音识别与说话人验证分析揭示了如Whisper等专有模型的局限性,强调开放数据模型对可靠儿童语音研究的重要性。所有实验均基于ESPnet,公开基准为构建鲁棒儿童语音处理系统提供关键洞见。
原文摘要 · Abstract (English)
Despite advancements in ASR, child speech recognition remains challenging due to acoustic variability and limited annotated data. While fine-tuning adult ASR models on child speech is common, comparisons with flat-start training remain underexplored. We compare flat-start training across multiple datasets, SSL representations (WavLM, XEUS), and decoder architectures. Our results show that SSL representations are biased toward adult speech, with flat-start training on child speech mitigating these biases. We also analyze model scaling, finding consistent improvements up to 1B parameters, beyond which performance plateaus. Additionally, age-related ASR and speaker verification analysis highlights the limitations of proprietary models like Whisper, emphasizing the need for open-data models for reliable child speech research. All investigations are conducted using ESPnet, and our publicly available benchmark provides insights into training strategies for robust child speech processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。