用加权交叉熵提升低资源语言在多语言语音识别中的表现
Weighted Cross-entropy for Low-Resource Languages in Multilingual Speech Recognition
- 通过动态加权交叉熵,让低资源语言在训练中获得更多关注
- 低资源语言WER降低6.69%,相比未优化模型减少48.86%
- 六种语言平均WER降3.29%,高资源语言无性能下降
本文针对将低资源语言融入多语言自动语音识别(ASR)系统所面临的挑战,提出一种加权交叉熵的新应用——该方法通常用于处理数据分布不均的情况。在持续多语言学习背景下,我们对Whisper多语言ASR模型在五个高资源语言和一个低资源语言上进行微调,采用语言加权的动态交叉熵与数据增强策略。实验结果显示,相较于未使用本方法的微调模型,低资源语言的词错误率(WER)降低了6.69%;相比原始Whisper模型,更是减少了48.86%。此外,该方法在六种语言上实现了平均3.29%的WER降低,且高资源语言未出现性能退化。
原文摘要 · Abstract (English)
This paper addresses the challenge of integrating low-resource languages into multilingual automatic speech recognition (ASR) systems. We introduce a novel application of weighted cross-entropy, typically used for unbalanced datasets, to facilitate the integration of low-resource languages into pre-trained multilingual ASR models within the context of continual multilingual learning. We fine-tune the Whisper multilingual ASR model on five high-resource languages and one low-resource language, employing language-weighted dynamic cross-entropy and data augmentation. The results show a remarkable 6.69% word error rate (WER) reduction for the low-resource language compared to the fine-tuned model without applying our approach, and a 48.86% WER reduction compared to the original Whisper model. In addition, our approach yields an average WER reduction of 3.29% across the six languages, showing no degradation for the high-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。