arXiv:2508.21225eess.AScs.AI2025-08中稿 · ed被引 4

用预训练模型的层间特征提升儿童语音零样本识别准确率

Can Layer-wise SSL Features Improve Zero-Shot ASR Performance for Children's Speech?

  • 从Wav2Vec2等模型提取分层特征,融入DNN声学模型
  • Wav2Vec2第22层使儿童语音WER降至5.15%,相对降低51.64%
  • 对不同年龄组儿童均有效,尤其适合低龄儿童语音识别

自动语音识别(ASR)系统在处理儿童语音时表现不佳,因其声学与语言特征差异大且变化显著。尽管自监督学习(SSL)模型显著提升了成人语音识别性能,儿童语音识别仍是难题。本研究评估了Wav2Vec2、HuBERT、Data2Vec和WavLM等先进SSL模型的层间特征,在零样本场景下改进儿童语音识别效果。通过Kaldi工具包将特征集成至简化的DNN声学模型中。分析发现,使用成人语料WSJCAM0训练、儿童语料PFSTAR测试时,Wav2Vec2第22层特征使词错误率(WER)降至5.15%,相比直接零样本解码(WER 10.65%)相对降低51.64%。年龄分组分析显示,随年龄增长性能持续提升,即使低龄组也有显著改善。在CMU Kids数据集上的实验验证了该方法的通用性。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) systems often struggle to accurately process children's speech due to its distinct and highly variable acoustic and linguistic characteristics. While recent advancements in self-supervised learning (SSL) models have greatly enhanced the transcription of adult speech, accurately transcribing children's speech remains a significant challenge. This study investigates the effectiveness of layer-wise features extracted from state-of-the-art SSL pre-trained models - specifically, Wav2Vec2, HuBERT, Data2Vec, and WavLM in improving the performance of ASR for children's speech in zero-shot scenarios. A detailed analysis of features extracted from these models was conducted, integrating them into a simplified DNN-based ASR system using the Kaldi toolkit. The analysis identified the most effective layers for enhancing ASR performance on children's speech in a zero-shot scenario, where WSJCAM0 adult speech was used for training and PFSTAR children speech for testing. Experimental results indicated that Layer 22 of the Wav2Vec2 model achieved the lowest Word Error Rate (WER) of 5.15%, representing a 51.64% relative improvement over the direct zero-shot decoding using Wav2Vec2 (WER of 10.65%). Additionally, age group-wise analysis demonstrated consistent performance improvements with increasing age, along with significant gains observed even in younger age groups using the SSL features. Further experiments on the CMU Kids dataset confirmed similar trends, highlighting the generalizability of the proposed approach.

语音识别自监督学习儿童语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。