arXiv:2606.19793eess.AScs.AI2026-06被引 3

针对口吃语音识别难题,优化特征与模型组合提升识别准确率。

Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models

论文配图:Systematic Study of Dysarthric Speech Recognition: Spectral Features and Acoustic Models
图 1 · 摘自论文原文
  • 针对口吃语音设计多组声学特征与模型组合方案。
  • 在隔离词和句子识别上分别提升4.65%和4.63%相对性能。
  • 适合语音识别、残障辅助技术研究者参考。

口吃语音识别的主要挑战源于发音不精确带来的显著声学差异。以往研究显示,采用混合DNN/HMM序列判别训练可提升识别效果。本文系统考察了多种声学特征与不同声学模型的组合,为各类模型提供合适特征选择。引入基频(Pitch)特征显著提升了识别性能,尤其在口吃语音的句子识别任务中表现突出。通过对TORGO数据库的系统分析,验证了因子化时间延迟神经网络(F-TDNN)在口吃语音识别中的潜力。所提方法在F-TDNN模型上实现孤立词识别4.65%、句子识别4.63%的相对性能提升,有效缓解了语音变异影响,归因于对连续训练样本块间重叠帧数的精心设计。

原文摘要 · Abstract (English)

The challenge associated with recognizing dysarthric speech primarily arises from pronounced acoustic variability attributed to impaired articulatory precision. Past research has demonstrated improved recognition through the use of hybrid DNN/HMM sequence discriminative training. This paper presents a comprehensive investigation of various combinations of acoustic features tailored to different Acoustic Models, offering suitable feature selections for each. The incorporation of Pitch features notably improved recognition performance, especially for sentence recognition tasks involving dysarthric speech. Through a systematic examination of the TORGO database, we have demonstrated the potential to enhance the performance of the state-of-the-art Factorized Time Delay Neural Network (F-TDNN) model for recognizing dysarthric speech. Our methods, implemented with the F-TDNN model, resulted in a 4.65\% relative improvement in isolated word recognition and a 4.63\% relative improvement in sentence recognition for dysarthric speech, compared to previous research. This improvement effectively compensates for speech variability, attributable to our deliberate selection of the number of overlapping frames between consecutive training example chunks.

语音识别口吃语音声学模型特征优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。