评估大模型在低资源窄带语音上的表现并提出改进方案
Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings
- 在印度语和印式英语上测试主流ASR模型性能
- 微调可提升效果但受预训练数据量影响显著
- 揭示了窄带+低资源场景下的关键挑战
全球电话通话多通过窄带信道进行,且常为即兴口语。本文评估了广泛使用的开源与商用基础语音识别(ASR)模型在印地语(低资源语言)及印式英语(低资源口音)的窄带通话中的表现。首先在零样本设置下测试,发现所有模型性能均不理想。进一步研究使用有限真实录音微调开源模型的效果,结果表明微调虽带来一定提升,但成效因语言和口音而异,主要受预训练阶段数据量影响。
原文摘要 · Abstract (English)
Telephony conversations worldwide are conducted over narrow-band channels and are often spontaneous and colloquial in nature. This paper evaluates the performance of widely used foundational automatic speech recognition (ASR) models -- both open-source and commercial -- on narrow-band conversations in Hindi, a low-resource language, and Indian-accented English, a low-resource accent. We first assess these models in a zero-shot setting and find that their performance remains suboptimal across the board. Highlighting the challenges faced by ASR models in narrow-band and low-resource language scenarios, we further investigate the impact of fine-tuning open-source models using a limited set of real-life annotated recordings. Our findings indicate that while fine-tuning provides some improvements, its effectiveness varies across languages and accents, largely influenced by the amount of data encountered during pretraining
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。