arXiv:2410.00940cs.CL2024-10被引 2

用少量数据微调多语言模型,实现对低资源语言伊卡语的高精度语音识别

Automatic Speech Recognition for the Ika Language

  • 在1小时高质量数据上微调wav2vec 2.0多语言模型
  • 达到WER 0.5377、CER 0.2651的低错误率
  • 适合资源匮乏语言的语音识别研究者参考

我们提出一种经济高效的方案,用于开发低资源语言伊卡语的自动语音识别(ASR)模型。通过在源自伊卡语新约圣经翻译的高质量语音数据集上微调预训练的wav2vec 2.0大规模多语言语音模型,实验表明,仅需超过1小时的训练数据即可实现词错误率(WER)0.5377、字符错误率(CER)0.2651。参数量达10亿的大模型因具备更强表达能力而优于3亿参数的小模型。然而,由于训练数据量小,模型出现过拟合,影响泛化性能。研究证明了利用多语言预训练模型在低资源语言上的潜力。未来工作应聚焦于扩充数据集并探索缓解过拟合的技术。

原文摘要 · Abstract (English)

We present a cost-effective approach for developing Automatic Speech Recognition (ASR) models for low-resource languages like Ika. We fine-tune the pretrained wav2vec 2.0 Massively Multilingual Speech Models on a high-quality speech dataset compiled from New Testament Bible translations in Ika. Our results show that fine-tuning multilingual pretrained models achieves a Word Error Rate (WER) of 0.5377 and Character Error Rate (CER) of 0.2651 with just over 1 hour of training data. The larger 1 billion parameter model outperforms the smaller 300 million parameter model due to its greater complexity and ability to store richer speech representations. However, we observe overfitting to the small training dataset, reducing generalizability. Our findings demonstrate the potential of leveraging multilingual pretrained models for low-resource languages. Future work should focus on expanding the dataset and exploring techniques to mitigate overfitting.

语音识别低资源语言wav2vec2多语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。