arXiv:2608.19361cs.CLeess.AS2026-08

为低资源语言米佐语构建语音识别系统并验证模型迁移效果

A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation

论文配图:A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
图 1 · 摘自论文原文
  • 用17.62小时数据微调Whisper和SraVaani模型
  • 形态感知评估下最低词错误率降至7.22%
  • 适合低资源语言语音识别研究者参考

本研究针对米佐语这一低资源语言,构建了自动语音识别(ASR)系统。通过收集17.62小时的语音数据并进行清洗,对三种Whisper多语言模型及一个SraVaani 1.0印地语多语言模型进行微调。Whisper-large-v3在常规评估中达到最低词错误率(WER)18.08%,在形态感知评估下进一步降至7.22%。SraVaani 1.0零样本评估的WER为58.27%,经米佐语特化微调后,常规WER降至29.45%,形态感知WER降至17.93%。结果表明,即使在未见语言上,Whisper模型也能实现显著低错率;而SraVaani 1.0虽已包含米佐语支持,但使用精心标注的数据微调可大幅提升性能。

原文摘要 · Abstract (English)

This study reports the development of an Automatic Speech Recognition (ASR) system in Mizo, a low-resource language. The development included collecting 17.62 hours of speech data, curating it, and fine-tuning the Mizo ASR system with three Whisper multilingual models and with the SraVaani 1.0 Indic multilingual model. Whisper-large-v3 achieved the lowest conventional WER (18.08%), while morphology-aware evaluation yielded a WER of 7.22%. Zero-shot evaluation of the SraVaani 1.0 Indic multilingual model yielded a WER of 58.27%, while Mizo-specific fine-tuning reduced the conventional WER to 29.45% and the morphology-aware WER to 17.93%. The results demonstrate that the Whisper model can achieve a substantially low WER, even when adapted to an unseen language. In contrast, SraVaani 1.0 supports the Mizo language in its multilingual model; however, fine-tuning with carefully curated Mizo speech data substantially improves its performance.

语音识别低资源语言Whisper形态感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。