通过无监督节奏与音色转换,提升失语症语音的识别准确率。
Unsupervised Rhythm and Voice Conversion to Improve ASR on Dysarthric Speech
- 基于音节的节奏建模,适配失语症语音特点。
- LF-MMI模型在严重失语病例上词错误率显著降低。
- 适合需要提升失语症语音识别的无障碍应用开发。
自动语音识别(ASR)系统因失语症语音存在高说话人差异性和慢语速而表现不佳。为此,本文探索将失语症语音转换为正常语音以提升识别性能。方法扩展了节奏与音色(RnV)转换框架,引入适用于失语症语音的音节级节奏建模。通过训练LF-MMI模型并微调Whisper在转换后的语音上评估效果。在Torgo语料库上的实验表明,LF-MMI模型在严重失语病例中词错误率显著下降,而微调Whisper在转换数据上的性能提升有限。结果表明,无监督节奏与音色转换对失语症语音识别具有潜力。代码已公开:https://github.com/idiap/RnV
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our approach extends the Rhythm and Voice (RnV) conversion framework by introducing a syllable-based rhythm modeling method suited for dysarthric speech. We assess its impact on ASR by training LF-MMI models and fine-tuning Whisper on converted speech. Experiments on the Torgo corpus reveal that LF-MMI achieves significant word error rate reductions, especially for more severe cases of dysarthria, while fine-tuning Whisper on converted data has minimal effect on its performance. These results highlight the potential of unsupervised rhythm and voice conversion for dysarthric ASR. Code available at: https://github.com/idiap/RnV
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。