arXiv:2503.23542cs.CL2025-03被引 10

用语言模型提升小语种语音识别准确率,效果显著。

Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages

  • 将语言模型与微调后的Whisper结合,增强小语种识别能力。
  • 在低资源场景下词错误率最高降低51%,外分布句子降34%。
  • 适合关注多语言语音识别与小语种技术落地的研究者。

自动语音识别系统在融合多语言、多任务模型(如Whisper)后取得显著进展,具备跨语言理解与处理能力。然而,这些模型在少数语言的语义差异处理上仍显不足。本研究通过将传统及新型语言模型与微调后的Whisper模型结合,提升其在非主流语言中的表现。在多个数据集上进行严格微调与评估,结果显示词错误率显著下降,尤其在低资源场景中:使用统计语言模型时,分布内数据集最高降低51%,分布外句子降低34%;大语言模型则在多样语言环境中提供稳定但适度的改进。结果表明,该方法对所有规模模型均有效,但提升幅度取决于语言模型参数优化程度。研究强调报告基于Transformer的语音识别结果时需注意评估参数选择。本工作为实现更包容的语音识别技术铺平道路,通过丰富语言知识提升跨语言性能。更多实现细节见http://www.github.com/hitz-zentroa/whisper-lm。

原文摘要 · Abstract (English)

Automatic speech recognition systems have undoubtedly advanced with the integration of multilingual and multitask models such as Whisper, which have shown a promising ability to understand and process speech across a wide range of languages. Despite their robustness, these models often fall short in handling the linguistic distinctions of minority languages. This study addresses this gap by integrating traditional and novel language models with fine-tuned Whisper models to raise their performance in less commonly studied languages. Through rigorous fine-tuning and evaluation across multiple datasets, we demonstrate substantial improvements in word error rate, particularly in low-resource scenarios. Our approach not only does take advantage of the extensive data Whisper was pre-trained on, but also complements its linguistic adaptability by incorporating language models. We obtained improvements up to 51% for in-distribution datasets and up to 34% for out-of-distribution sentences using statistical language models, while large language models provided moderate but consistently robust improvement across diverse linguistic contexts. The findings reveal that, while the integration reliably benefits all model sizes, the extent of improvement varies, highlighting the importance of optimized language model parameters. Finally, we emphasize the importance of selecting appropriate evaluation parameters when reporting the results using transformer-based ASR models. In summary, this research clears the way for more inclusive ASR technologies that perform better across languages by enriching their linguistic knowledge. For further implementation details of this study, the technical documentation and source code are available at http://www.github.com/hitz-zentroa/whisper-lm.

语音识别小语种语言模型Whisper

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。