用少量失语语音数据微调,显著提升语音识别公平性。
Towards a Single ASR Model That Generalizes to Disordered Speech
- 用约1000小时失语语音微调主流语音模型。
- 失语语音识别准确率提升33%(提示语音)和26%(自然对话)。
- 不牺牲标准语音识别性能,适合关注无障碍技术的开发者。
本研究探讨将约1000小时失语语音数据融入近顶尖语音识别(ASR)基线系统的微调过程。尽管该数据不足整体训练数据的1%,但实验显示在失语语音识别上取得显著提升:提示语音上准确率提高33%,新收集的自发对话数据集上提高26%。更重要的是,标准语音识别基准性能无明显下降。此外,该微调策略使基线系统与个性化模型之间的差距缩小了64%,表明进步显著但仍存优化空间。从公平性角度出发,此结果表明,在训练流程中加入少量高质量失语语音数据,是提升语音技术对言语障碍用户可及性的简单有效路径。
原文摘要 · Abstract (English)
This study investigates the impact of integrating a dataset of disordered speech recordings ($\sim$1,000 hours) into the fine-tuning of a near state-of-the-art ASR baseline system. Contrary to what one might expect, despite the data being less than 1% of the training data of the ASR system, we find a considerable improvement in disordered speech recognition accuracy. Specifically, we observe a 33% improvement on prompted speech, and a 26% improvement on a newly gathered spontaneous, conversational dataset of disordered speech. Importantly, there is no significant performance decline on standard speech recognition benchmarks. Further, we observe that the proposed tuning strategy helps close the gap between the baseline system and personalized models by 64% highlighting the significant progress as well as the room for improvement. Given the substantial benefits of our findings, this experiment suggests that from a fairness perspective, incorporating a small fraction of high quality disordered speech data in a training recipe is an easy step that could be done to make speech technology more accessible for users with speech disabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。