用少量个性化数据微调大模型,显著提升失语症语音识别准确率。
Adapting Foundation ASR Models to Dysarthric Speech: A Case Study

- 基于Whisper模型,仅用1.4小时语音数据微调即达15.8%错误率。
- 使用全部92+8.8小时数据后,读音测试错误率降至9.7%,纠错测试降为7.8%。
- 证明个性化微调对失语症语音识别实用有效,适合真实场景部署。
自动语音识别(ASR)系统在失语症语音上表现不佳,限制了其对患者日常交流的帮助。本文针对一位失语症说话者构建个性化ASR系统,通过将基础ASR模型适配到特定说话者数据。利用TEQST工具,我们收集了92小时朗读语音,并通过部署的移动应用获取了8.8小时用户修正数据。从Whisper开始微调,在仅1.4小时适配数据下,读音测试集词错误率降至15.8%;使用22.5小时数据时,读音与修正测试集分别达到10.7%和16.1%;使用全部数据后,读音测试集最优达9.7%,修正测试集达7.8%。使用LoRA或Qwen3-ASR作为基础模型在此任务中表现更差。结果表明,个性化微调可大幅提高基础模型在失语症语音上的有效性,具备实际部署潜力。
原文摘要 · Abstract (English)
Automatic speech recognition (ASR) systems often perform poorly in dysarthric speech, limiting their usefulness to affected speakers in everyday communication. This paper presents a personalized ASR system for a dysarthric speaker, built by adapting a foundation ASR model to speaker-specific data. Using the TEQST tool, we collected 92 hours of read speech and later added 8.8 hours of user corrections gathered through a deployed mobile application. Starting from Whisper, fine-tuning reduced word error rate to 15.8% with only 1.4 hours of adaptation data on the read test set, reached 10.7% / 16.1% with 22.5 hours on the read and corrections test sets, respectively, and achieved the best result of 9.7% when using all available data including the corrections on the read test set and 7.8% on the corrections test set. Using LoRA adaptation and/or Qwen3-ASR as foundation model performed worse in this setting. The results show that personalized fine-tuning can make foundation ASR models substantially more effective for dysarthric speech and suitable for practical deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。