arXiv:2409.19818eess.AScs.SD2024-09被引 18

用帕金森患者语音微调语音识别模型,显著提升识别准确率。

Fine-Tuning Automatic Speech Recognition for People with Parkinson's: An Effective Strategy for Enhancing Speech Technology Accessibility

  • 多任务学习同时识别语音并预测患者病情严重程度
  • 相比仅用LibriSpeech数据微调,错误率降低超26%
  • 适合关注语音技术无障碍应用的研究者与开发者

本文通过在2023-10-05版语音可及性项目(SAP)数据包上微调预训练自动语音识别(ASR)模型,提升帕金森病患者的构音障碍和声音异常语音识别效果。该数据包包含253名帕金森患者语音。实验测试了对脑瘫有效的多种方法,包括说话人聚类、严重程度依赖模型、加权微调和多任务学习。最佳结果来自多任务学习模型,该模型在输出语音转录的同时,额外预测说话人病情严重程度。相比仅使用Librispeech数据微调的基线模型,该方法在100小时和960小时LibriSpeech数据微调基础上,分别实现37.62%和26.97%的词错误率降低。

原文摘要 · Abstract (English)

This paper enhances dysarthric and dysphonic speech recognition by fine-tuning pretrained automatic speech recognition (ASR) models on the 2023-10-05 data package of the Speech Accessibility Project (SAP), which contains the speech of 253 people with Parkinson's disease. Experiments tested methods that have been effective for Cerebral Palsy, including the use of speaker clustering and severity-dependent models, weighted fine-tuning, and multi-task learning. Best results were obtained using a multi-task learning model, in which the ASR is trained to produce an estimate of the speaker's impairment severity as an auxiliary output. The resulting word error rates are considerably improved relative to a baseline model fine-tuned using only Librispeech data, with word error rate improvements of 37.62\% and 26.97\% compared to fine-tuning on 100h and 960h of LibriSpeech data, respectively.

语音识别帕金森无障碍技术多任务学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。