arXiv:2501.14994cs.SDcs.AI2025-01中稿 · ICASSP 2025被引 8

构建可泛化到不同病因和说话人的发音障碍语音识别模型

Robust Cross-Etiology and Speaker-Independent Dysarthric Speech Recognition

  • 基于Whisper框架设计无特定说话人依赖的识别系统
  • 在帕金森病数据集上达6.99%词错误率,跨病因测试中39.56%词错误率
  • 首次验证模型对不同病因发音障碍的通用性,适合无障碍语音交互研究

本文提出一种无特定说话人依赖的发音障碍语音识别系统,重点评估新发布的Speech Accessibility Project (SAP-1005)数据集,该数据集包含帕金森病(PD)患者的语音样本。尽管现有研究多为说话人依赖型,限制了跨说话人与病因的泛化能力,本研究旨在开发鲁棒的无特定说话人模型,实现对不同发音障碍类型的一致识别。作为次要目标,进一步在含脑瘫(CP)和肌萎缩侧索硬化症(ALS)患者的TORGO数据集上测试模型的跨病因性能。基于Whisper模型,系统在SAP-1005上达到6.99%的字符错误率(CER)和10.71%的词错误率(WER);在跨病因测试中,于TORGO数据集取得25.08%的CER和39.56%的WER。结果表明该方法具备良好的跨说话人与跨病因泛化能力。

原文摘要 · Abstract (English)

In this paper, we present a speaker-independent dysarthric speech recognition system, with a focus on evaluating the recently released Speech Accessibility Project (SAP-1005) dataset, which includes speech data from individuals with Parkinson's disease (PD). Despite the growing body of research in dysarthric speech recognition, many existing systems are speaker-dependent and adaptive, limiting their generalizability across different speakers and etiologies. Our primary objective is to develop a robust speaker-independent model capable of accurately recognizing dysarthric speech, irrespective of the speaker. Additionally, as a secondary objective, we aim to test the cross-etiology performance of our model by evaluating it on the TORGO dataset, which contains speech samples from individuals with cerebral palsy (CP) and amyotrophic lateral sclerosis (ALS). By leveraging the Whisper model, our speaker-independent system achieved a CER of 6.99% and a WER of 10.71% on the SAP-1005 dataset. Further, in cross-etiology settings, we achieved a CER of 25.08% and a WER of 39.56% on the TORGO dataset. These results highlight the potential of our approach to generalize across unseen speakers and different etiologies of dysarthria.

语音识别发音障碍跨病因Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。