用数据增强和大模型纠错提升口腔癌患者语音识别准确率
Improving Automatic Speech Recognition for Speakers Treated for Oral Cancer using Data Augmentation and LLM Error Correction
- 通过文本转语音生成合成数据,扩充稀缺的口腔癌语音样本
- 结合大模型纠错使识别错误率降低21.4%至26.2%(微调后)
- 对语音障碍人群有实用价值,适合医疗辅助与无障碍技术研究
近年来,自动语音识别(ASR)系统性能显著提升,但针对口腔癌(OC)康复患者的语音识别仍表现不佳。由于此类语音数据稀缺且差异大,模型训练困难。本文在荷兰口腔癌语音语料库上应用多种数据增强技术生成合成数据,并评估其对ASR性能的影响。对Whisper和多语言语音模型(MMS)进行微调,发现使用文本到语音(TTS)生成数据后,平均词错误率(WER)相对降低8%。结合大语言模型(LLM)纠错,微调模型的WER进一步降低21.4%-26.2%,非微调模型降低10.0%。最终,Whisper实现40%相对WER下降,MMS达50%,表明数据增强与LLM纠错结合是提升口腔癌语音识别的有效策略。
原文摘要 · Abstract (English)
In recent years, the performance of automatic speech recognition (ASR) systems has made considerable progress. Unfortunately, for people with speech impairments, such as people treated for oral cancer (OC), ASR performance is still lagging behind. The scarcity and variability of OC speech data makes development of ASR models for this type of speech difficult. In this work, we use data augmentation and large language model (LLM) error correction to mitigate this problem. We apply various augmentation techniques on a corpus of Dutch oral cancer speech to create synthetic data, and evaluate their effect on ASR performance. We finetune Whisper and Massively Multilingual Speech (MMS) models for each augmentation technique and observe, on average, an 8% relative decrease in Word Error Rate (WER) when including data created using text-to-speech (TTS). When employing LLMs for error correction, we see a further 21.4-26.2% relative decrease in WER for finetuned ASR models and a 10.0% relative decrease for non-finetuned models. Overall, we achieve a 40% relative WER decrease for Whisper and a 50% relative WER decrease for MMS, indicating that a combination of data augmentation and LLM correction is a viable strategy for the recognition of OC speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。