arXiv:2503.18485cs.CLcs.LG2025-03被引 10

微调Whisper提升阿姆哈拉语识别准确率

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

  • 用共用语音、FLEURS和BDU-speech数据微调Whisper模型
  • 混合新旧数据使词错误率显著降低,同音词归一化提升性能
  • 适合低资源语言语音识别研究者参考

本研究探索了对OpenAI的Whisper自动语音识别(ASR)模型进行微调,以提升阿姆哈拉语(一种低资源语言)的转录准确率。尽管基础Whisper模型因训练数据中阿姆哈拉语代表性不足而表现不佳,但通过使用Mozilla Common Voice、FLEURS及BDU-speech等数据集进行微调,取得了明显改进。最佳模型Whispersmall-am在结合现有FLEURS数据与新未见阿姆哈拉语数据时表现最优。仅使用新数据训练导致性能下降,而融合FLEURS数据则强化模型,使其更好地专精于阿姆哈拉语。此外,对阿姆哈拉语同音词进行归一化处理,显著提升了词错误率(WER)和双语评估替代(BLEU)分数。该研究强调了微调策略与数据组合对低资源语言语音识别的重要性,为未来阿姆哈拉语语音识别研究提供重要启示。

原文摘要 · Abstract (English)

This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic, a low-resource language, to improve transcription accuracy. While the foundational Whisper model struggles with Amharic due to limited representation in its training data, we fine-tune it using datasets like Mozilla Common Voice, FLEURS, and the BDU-speech dataset. The best-performing model, Whispersmall-am, significantly improves when finetuned on a mix of existing FLEURS data and new, unseen Amharic datasets. Training solely on new data leads to poor performance, but combining it with FLEURS data reinforces the model, enabling better specialization in Amharic. We also demonstrate that normalizing Amharic homophones significantly enhances Word Error Rate (WER) and Bilingual Evaluation Understudy (BLEU) scores. This study underscores the importance of fine-tuning strategies and dataset composition for improving ASR in low-resource languages, providing insights for future Amharic speech recognition research.

语音识别低资源语言Whisper阿姆哈拉语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。