arXiv:2501.10256eess.AScs.AI2025-01中稿 · ICASSP 2025 Satell…被引 7

无需标注数据,将失语语音转换为自然语音提升识别率

Unsupervised Rhythm and Voice Conversion of Dysarthric to Healthy Speech for ASR

  • 用自监督表征实现无监督节奏与音色转换
  • 对严重失语者语音识别准确率显著提升
  • 适合无障碍语音识别与残障辅助技术研究者

自动语音识别系统在失语语音上的表现不佳。以往方法通过调整语速来减少与正常语音的差异,但需依赖带标注的语音数据估计语速和音素时长,对未见说话人不适用。为此,本文基于自监督语音表示,结合无监督节奏与音色转换方法,将失语语音映射为典型语音。我们使用在健康语音上预训练的大规模ASR模型(未进一步微调)评估输出效果,发现该节奏转换方法对Torgo语料库中重度失语者尤为有效。代码与音频样例已公开于 https://idiap.github.io/RnV。

原文摘要 · Abstract (English)

Automatic speech recognition (ASR) systems are well known to perform poorly on dysarthric speech. Previous works have addressed this by speaking rate modification to reduce the mismatch with typical speech. Unfortunately, these approaches rely on transcribed speech data to estimate speaking rates and phoneme durations, which might not be available for unseen speakers. Therefore, we combine unsupervised rhythm and voice conversion methods based on self-supervised speech representations to map dysarthric to typical speech. We evaluate the outputs with a large ASR model pre-trained on healthy speech without further fine-tuning and find that the proposed rhythm conversion especially improves performance for speakers of the Torgo corpus with more severe cases of dysarthria. Code and audio samples are available at https://idiap.github.io/RnV .

语音转换失语识别自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。