研究wav2vec2模型跨语言迁移能力,发现语言相似性比训练数据量更重要。
On the Cross-lingual Transferability of Pre-trained wav2vec2-based Models
- 用15个预训练模型在18种语言上测试语音识别性能,分析跨语言迁移规律。
- 语言多样性比训练数据量更影响模型表现,印欧语系语言效果优于非印欧语系。
- 单语模型也能实现正向跨语言迁移,尤其当预训练语言与目标语言相近时。
使用大型预训练模型提供的表征已成为在多种任务中取得最先进结果的主要策略。近期提出的wav2vec 2.0模型对语音数据的大型预训练工作具有开创性意义。许多模型采用与wav2vec 2.0相同的架构进行预训练,并在各类语音相关任务中达到最先进水平。先前研究已表明,预训练阶段所用数据会影响模型在下游任务中的表现,应在使用前予以考虑。然而,很少有研究进一步探讨这些wav2vec2模型在不同语言间的知识迁移行为,尤其是目标语言与预训练语言不同时的情况。本研究旨在探究wav2vec2模型的跨语言可迁移性。我们在18种语言上对15个大型预训练模型进行了多个微调实验,评估其在语音识别任务中的表现。结果显示,预训练数据规模对最终性能的影响不如语言多样性重要。我们观察到,印欧语系语言的表现优于非印欧语系语言。此外,即使使用单语模型,也观察到显著的正向跨语言知识迁移现象,且在预训练语言与下游任务语言更相似时更为明显。基于这些发现,我们希望帮助科研界更好地利用现有wav2vec2模型,并指导新模型的预训练工作。
原文摘要 · Abstract (English)
Using representations provided by a large pre-trained model has become the primary strategy for achieving state-of-the-art results in a wide range of tasks. A recently proposed large pre-trained model, wav2vec 2.0, was seminal for several other works on pre-training large models on speech data. Many models are being pre-trained using the same architecture as wav2vec 2.0 and are getting state-of-the-art in various speech-related tasks. Previous work has demonstrated that the data used during the pre-training of these wav2vec2-based models can impact the model's performance in downstream tasks, and this should be taken into consideration before utilizing these models. However, few works have proposed investigating further how the transfer knowledge of these pre-trained models behaves in different languages, even when the target language differs from the one used during the model's pre-training. Our work aims to investigate the cross-lingual transferability of these wav2vec2-based models. We performed several fine-tuning experiments on the speech recognition task in 18 languages using 15 large pre-trained models. The results of our experiments showed us that the size of data used during the pre-training of these models is not as important to the final performance as the diversity. We noticed that the performance of Indo-European languages is superior to non-Indo-European languages in the evaluated models. We have observed a positive cross-lingual transfer of knowledge using monolingual models, which was evident in all the languages we used, but more pronounced when the language used during pre-training was more similar to the downstream task language. With these findings, we aim to assist the scientific community in utilizing existing wav2vec2-based pre-trained models, as well as facilitate the pre-training of new ones.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。