优化特征选择与训练块重叠,提升罕见儿童语音识别准确率
Cross-Dataset, Age, and Gender Generalization: A Comprehensive Analysis of Fine-Tuning Strategies for Low-Resource Children's ASR

- 针对发音障碍语音设计特定声学特征组合与训练块重叠策略
- 在孤立词和句子识别上分别提升4.65%和4.63%相对性能
- 适用于低资源儿童语音识别,尤其对发音不清晰者有效
发音障碍语音识别的主要挑战源于发音器官控制能力受损带来的显著声学变异。以往研究已证明,采用混合DNN/HMM序列判别训练可改善识别效果。本文系统性地探讨了针对不同声学模型的声学特征组合,提出了适合各类模型的特征选择方案。引入音高特征显著提升了识别性能,尤其在发音障碍者的句子识别任务中表现突出。通过对TORGO数据库的深入分析,验证了在因子化时延神经网络(F-TDNN)模型上应用本方法的潜力。实验结果显示,相较于先前研究,该方法在孤立词识别上实现4.65%的相对性能提升,在句子识别上实现4.63%的相对提升。这一改进有效缓解了因发音不精准导致的语音变异性问题,归因于对连续训练样本块之间重叠帧数的精心设计。
原文摘要 · Abstract (English)
The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。