arXiv:2410.16330eess.AScs.CL2024-10被引 10

用预训练模型提升库尔德语语音识别准确率

End-to-End Transformer-based Automatic Speech Recognition for Northern Kurdish: A Pioneering Approach

  • 基于Whisper模型,采用模块化微调策略提升性能
  • 在68小时数据上实现WER 10.5%、CER 5.7%
  • 为低资源语言语音识别提供可复现的优化路径

低资源语言的自动语音识别(ASR)因训练数据有限而面临挑战。本文系统研究了预训练的Whisper模型在中东地区使用的北库尔德语(库尔曼吉语)上的有效性。我们对比了三种微调策略:基础微调、参数选择微调和新增模块微调。基于约68小时经验证的标注语音语料库,实验表明,采用新增模块的微调策略在专用测试集上显著提升识别准确率,使Whisper版本3的词错误率(WER)降至10.5%,字符错误率(CER)降至5.7%。结果证明,复杂Transformer模型在低资源场景下的潜力,强调了针对性微调技术对性能优化的关键作用。

原文摘要 · Abstract (English)

Automatic Speech Recognition (ASR) for low-resource languages remains a challenging task due to limited training data. This paper introduces a comprehensive study exploring the effectiveness of Whisper, a pre-trained ASR model, for Northern Kurdish (Kurmanji) an under-resourced language spoken in the Middle East. We investigate three fine-tuning strategies: vanilla, specific parameters, and additional modules. Using a Northern Kurdish fine-tuning speech corpus containing approximately 68 hours of validated transcribed data, our experiments demonstrate that the additional module fine-tuning strategy significantly improves ASR accuracy on a specialized test set, achieving a Word Error Rate (WER) of 10.5% and Character Error Rate (CER) of 5.7% with Whisper version 3. These results underscore the potential of sophisticated transformer models for low-resource ASR and emphasize the importance of tailored fine-tuning techniques for optimal performance.

语音识别低资源语言WhisperTransformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。