为低资源语言米佐语构建语音识别系统并验证模型迁移效果
A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

- 用17.62小时数据微调Whisper和SraVaani模型
- 形态感知评估下最低词错误率降至7.22%
- 适合低资源语言语音识别研究者参考
本研究针对米佐语这一低资源语言,构建了自动语音识别(ASR)系统。通过收集17.62小时的语音数据并进行清洗,对三种Whisper多语言模型及一个SraVaani 1.0印地语多语言模型进行微调。Whisper-large-v3在常规评估中达到最低词错误率(WER)18.08%,在形态感知评估下进一步降至7.22%。SraVaani 1.0零样本评估的WER为58.27%,经米佐语特化微调后,常规WER降至29.45%,形态感知WER降至17.93%。结果表明,即使在未见语言上,Whisper模型也能实现显著低错率;而SraVaani 1.0虽已包含米佐语支持,但使用精心标注的数据微调可大幅提升性能。
原文摘要 · Abstract (English)
This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。