用Whisper微调低资源原住民语言Baniwa的语音识别,首次实现37.5%的错误率。
Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
- 用1373段人工标注录音微调Whisper Small模型
- 在0.54小时语音上达到37.5% WER和7.45% CER
- 为极低资源原住民语言提供首个语音识别基线
近年来,大规模多语言基础模型使自动语音识别(ASR)取得显著进展,但多数成果仍集中于高资源语言,原住民语言仍面临语音资源匮乏和技术缺失问题。本文首次研究了将Whisper模型适配至巴西、哥伦比亚和委内瑞拉使用的原住民亚瓦堪语Baniwa的语音识别任务。实验基于1,373段人工转录的录音语料,总时长约0.54小时,以孤立词和短句为主。采用监督学习对Whisper Small模型进行微调,并通过词错误率(WER)和字符错误率(CER)评估。最佳模型达到37.5% WER和7.45% CER,证明多语言基础模型可成功应用于极端低资源原住民语言。该结果为Baniwa ASR建立了初始基准,为未来更大规模数据集、语言特定优化策略及后处理技术的研究奠定基础。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages, while indigenous languages continue to suffer from a lack of speech resources and language technologies. This work presents a preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela. The experiments were conducted using a corpus of 1,373 manually transcribed recordings obtained from a linguistic documentation project. The corpus contains approximately 0.54 hours of speech and consists primarily of isolated words and short elicited utterances. The Whisper Small model was fine-tuned using supervised learning and evaluated using Word Error Rate (WER) and Character Error Rate (CER). The best model achieved a WER of 37.5% and a CER of 7.45%, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages. The results establish an initial baseline for Baniwa Automatic Speech Recognition and provide a foundation for future research involving larger datasets, language-specific adaptation strategies, and post-processing techniques.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。