研究印地语方言间跨语言语音识别,发现小规模方言数据可媲美大规模标准语微调。
Dialect Matters: Cross-Lingual ASR Transfer for Low-Resource Indic Language Varieties
- 用真实嘈杂的混合语数据测试多种印地语方言间的迁移学习效果
- 小规模方言微调性能接近大规模标准语微调,方言距离非唯一决定因素
- 揭示预训练语言对非标准语音的偏见,适合低资源方言研究者参考
我们对大量印地语方言及语言变体中的自发性、嘈杂且混码的语音进行了跨语言迁移的实证研究。结果表明,尽管语言间亲缘关系越近,语音识别性能通常越高,但该因素无法完全解释方言场景下的表现差异。在多数情况下,使用较少的方言数据进行微调,其性能可与使用大量亲缘相关、高资源标准语数据微调相媲美。我们还以低资源的旁遮普方言——加尔瓦里语为例,评估了多个现代语音识别模型。最后,通过分析转录错误,考察了预训练语言对非标准语音的偏见,为方言和非标准化语音识别系统面临的挑战提供了进一步见解。
原文摘要 · Abstract (English)
We conduct an empirical study of cross-lingual transfer using spontaneous, noisy, and code-mixed speech across a wide range of Indic dialects and language varieties. Our results indicate that although ASR performance is generally improved with reduced phylogenetic distance between languages, this factor alone does not fully explain performance in dialectal settings. Often, fine-tuning on smaller amounts of dialectal data yields performance comparable to fine-tuning on larger amounts of phylogenetically-related, high-resource standardized languages. We also present a case study on Garhwali, a low-resource Pahari language variety, and evaluate multiple contemporary ASR models. Finally, we analyze transcription errors to examine bias toward pre-training languages, providing additional insight into challenges faced by ASR systems on dialectal and non-standardized speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。