用自建数据集微调Whisper,显著提升尼泊尔语语音识别准确率。
Whisper Finetuning on Nepali Language
- 构建包含多口音、方言和语境的尼泊尔语数据集并增强
- 小模型WER降36.2%,中模型降23.8%,优于原始Whisper
- 适合做低资源语言语音识别的开发者与研究者参考
尽管自动语音识别(ASR)模型持续进步,但对尼泊尔语等低资源语言的鲁棒模型开发仍具挑战。本研究构建了一个全面且通用的数据集,并对不同规模的OpenAI Whisper模型进行微调,以提升尼泊尔语语音转写准确率。我们结合公开ASR数据集与自录定制数据集,涵盖多样口音、方言与语调,并通过数据增强进一步丰富。实验表明,基于该数据集微调的模型在所有规模下均显著降低词错误率(WER),归因于说话人年龄、性别、情感、声学环境、方言差异及更密集音频片段(15-30秒)的多样性,以及人工校对的音视频内容。尤其,其性能优于在Fleur数据集上训练的Whisper基线模型,小模型最高降低36.2%,中模型降低23.8%。此外,数据增强有效提升了模型鲁棒性。结果强调了高质量、多样化与增强数据在适配前沿模型至低资源语言中的关键作用。
原文摘要 · Abstract (English)
Despite the growing advancements in Automatic Speech Recognition (ASR) models, the development of robust models for underrepresented languages, such as Nepali, remains a challenge. This research focuses on making an exhaustive and generalized dataset followed by fine-tuning OpenAI's Whisper models of different sizes to improve transcription (speech-to-text) accuracy for the Nepali language. We leverage publicly available ASR datasets and self-recorded custom datasets with a diverse range of accents, dialects, and speaking styles further enriched through augmentation. Our experimental results demonstrate that fine-tuning Whisper models on our curated custom dataset substantially reduces the Word Error Rate (WER) across all model sizes attributed to larger data variations in terms of speaker's age, gender, and sentiment, acoustic environment, dialect, denser audio segments (15-30 seconds) that are more compatible with Whisper's input, and manual curation of audios and transcriptions. Notably, our approach outperforms Whisper's baseline models trained on Fleur's dataset, achieving WER reductions of up to 36.2% on the small and 23.8% on medium models. Furthermore, we show that data augmentation plays a significant role in enhancing model robustness. Our approach underlines the importance of dataset quality, variation, and augmentation in the adaptation of state-of-the-art models to underrepresented languages for developing accurate ASR systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。