用多语言微调和语言识别提升低资源弗里斯兰语语音识别性能
Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance
- 通过多语言数据(弗里斯兰语、荷兰语、英语、德语)微调自监督模型
- 加入辅助语言识别任务后,弗里斯兰语识别准确率显著提升
- 方言语音识别效果差,数据采集方式影响结果,标准语评估可能高估实际表现
低资源语言的自动语音识别(ASR)性能仍远低于英语等高资源语言,主要因标注数据不足。当前先进方法采用自监督迁移学习,即在大量数据上预训练模型后,用少量目标语言标注数据进行微调。本文提出并评估了一种基于自监督学习的微调方法,用于提升弗里斯兰语及其地区方言(黏土弗里斯兰语、木弗里斯兰语、南弗里斯兰语)的识别性能。实验表明,使用包含弗里斯兰语、荷兰语、英语和德语的多语言微调数据,并引入辅助语言识别任务,可有效改善弗里斯兰语的识别表现。此外,研究发现方言语音识别性能明显下降,且这种差异受方言数据采集方式的影响。更重要的是,仅依赖标准语言数据评估可能导致对真实场景性能的高估,尤其在存在显著方言变异的语言中。
原文摘要 · Abstract (English)
Automatic Speech Recognition (ASR) performance for low-resource languages is still far behind that of higher-resource languages such as English, due to a lack of sufficient labeled data. State-of-the-art methods deploy self-supervised transfer learning where a model pre-trained on large amounts of data is fine-tuned using little labeled data in a target low-resource language. In this paper, we present and examine a method for fine-tuning an SSL-based model in order to improve the performance for Frisian and its regional dialects (Clay Frisian, Wood Frisian, and South Frisian). We show that Frisian ASR performance can be improved by using multilingual (Frisian, Dutch, English and German) fine-tuning data and an auxiliary language identification task. In addition, our findings show that performance on dialectal speech suffers substantially, and, importantly, that this effect is moderated by the elicitation approach used to collect the dialectal data. Our findings also particularly suggest that relying solely on standard language data for ASR evaluation may underestimate real-world performance, particularly in languages with substantial dialectal variation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。