用自监督模型提升法语儿童语音识别,尤其在嘈杂环境下表现更优。
Self-Supervised Models for Phoneme Recognition: Applications in Children's Speech for Reading Learning
- 用WavLM base+模型,通过解冻Transformer块微调,提升儿童语音识别效果。
- 在法语儿童语音数据上,该方法比基线模型性能显著提升。
- 模型在不同阅读任务和噪声水平下均更鲁棒,适合实际教学应用。
儿童语音识别仍属研究薄弱领域,主要受限于数据匮乏(尤其是非英语语言)及任务本身难度。本文基于前期工作,探索自监督模型在法语儿童语音识别中的应用。首先对比了wav2vec 2.0、HuBERT与WavLM模型在该任务的表现,选定性能最佳的WavLM base+进行后续实验。进一步通过微调时解冻其Transformer模块,显著提升模型性能,使其显著优于基础的Transformer+CTC模型。最后在真实应用场景下评估两模型行为,结果表明WavLM base+在多种阅读任务和噪声条件下更具鲁棒性。
原文摘要 · Abstract (English)
Child speech recognition is still an underdeveloped area of research due to the lack of data (especially on non-English languages) and the specific difficulties of this task. Having explored various architectures for child speech recognition in previous work, in this article we tackle recent self-supervised models. We first compare wav2vec 2.0, HuBERT and WavLM models adapted to phoneme recognition in French child speech, and continue our experiments with the best of them, WavLM base+. We then further adapt it by unfreezing its transformer blocks during fine-tuning on child speech, which greatly improves its performance and makes it significantly outperform our base model, a Transformer+CTC. Finally, we study in detail the behaviour of these two models under the real conditions of our application, and show that WavLM base+ is more robust to various reading tasks and noise levels. Index Terms: speech recognition, child speech, self-supervised learning
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。