微调Whisper提升痴呆语音识别准确率,还支持填充词检测
WhisperD: Dementia Speech Recognition and Filler Word Detection with Whisper
- 用痴呆语料微调Whisper,加入填充词识别
- 中等模型WER达0.24,优于现有方法
- 对未见数据和异常语音有良好泛化能力
Whisper因训练时未接触痴呆患者异常语音(如停顿、重复、句子碎片化),在转录痴呆语音时表现不佳。为实现低成本诊断与辅助技术开发,本文使用开源痴呆语音数据集DementiaBank及自建数据集对Whisper进行微调,显著降低词错误率(WER)。微调过程同时引入填充词识别,评估填充词包含率(FIR)与F1分数。结果表明,微调模型性能远超原始模型,中等规模模型达到WER 0.24,且在未见过的语音模式上仍具强泛化能力。
原文摘要 · Abstract (English)
Whisper fails to correctly transcribe dementia speech because persons with dementia (PwDs) often exhibit irregular speech patterns and disfluencies such as pauses, repetitions, and fragmented sentences. It was trained on standard speech and may have had little or no exposure to dementia-affected speech. However, correct transcription is vital for dementia speech for cost-effective diagnosis and the development of assistive technology. In this work, we fine-tune Whisper with the open-source dementia speech dataset (DementiaBank) and our in-house dataset to improve its word error rate (WER). The fine-tuning also includes filler words to ascertain the filler inclusion rate (FIR) and F1 score. The fine-tuned models significantly outperformed the off-the-shelf models. The medium-sized model achieved a WER of 0.24, outperforming previous work. Similarly, there was a notable generalisability to unseen data and speech patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。