用个性化语音合成提升严重构音障碍者的语音转写准确率
Improved Dysarthric Speech to Text Conversion via TTS Personalization
- 用患者未患病时的录音和声学嵌入插值生成可控严重程度的模拟构音障碍语音
- 结合真实与合成数据微调后,字符错误率从36-51%降至7.3%
- 适合残障辅助、语音康复及个性化语音识别研究者参考
我们针对一名匈牙利严重构音障碍患者开展了个案研究,构建定制化语音转文字系统。当前最先进的自动语音识别(ASR)模型在零样本情况下难以处理构音障碍语音,错误率高达36-51%。为缓解真实构音障碍数据稀缺问题,我们利用个性化文本转语音(TTS)系统生成合成语音,并对ASR模型进行微调。提出一种方法,通过融合患者病前录音与说话人嵌入插值,生成具有可控严重程度的合成构音障碍语音,实现不同损伤程度下的连续数据增强。在真实与合成数据上微调后,字符错误率(CER)从36-51%降至7.3%。我们的单语FastConformer_Hu ASR模型在相同数据上微调后显著优于Whisper-turbo,且合成语音带来18%的相对CER降低。结果表明,个性化ASR系统在提升重度言语障碍者可访问性方面潜力巨大。
原文摘要 · Abstract (English)
We present a case study on developing a customized speech-to-text system for a Hungarian speaker with severe dysarthria. State-of-the-art automatic speech recognition (ASR) models struggle with zero-shot transcription of dysarthric speech, yielding high error rates. To improve performance with limited real dysarthric data, we fine-tune an ASR model using synthetic speech generated via a personalized text-to-speech (TTS) system. We introduce a method for generating synthetic dysarthric speech with controlled severity by leveraging premorbidity recordings of the given speaker and speaker embedding interpolation, enabling ASR fine-tuning on a continuum of impairments. Fine-tuning on both real and synthetic dysarthric speech reduces the character error rate (CER) from 36-51% (zero-shot) to 7.3%. Our monolingual FastConformer_Hu ASR model significantly outperforms Whisper-turbo when fine-tuned on the same data, and the inclusion of synthetic speech contributes to an 18% relative CER reduction. These results highlight the potential of personalized ASR systems for improving accessibility for individuals with severe speech impairments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。