研究阿拉伯语模型在方言间的迁移能力,发现迁移效果与地理距离相关,且多方言训练可能产生负面干扰。
From FusHa to Folk: Exploring Cross-Lingual Transfer in Arabic Language Models
- 通过3个NLP任务和表征相似性分析方言迁移能力
- 方言间迁移效果不均,地理邻近方言迁移更优
- 多方言联合训练反而导致性能下降,提示方言差异不可忽视
阿拉伯语语言模型主要在现代标准阿拉伯语(MSA)上预训练,预期能迁移到各地域方言。尽管MSA用于正式场合,人们在线交流时使用多种方言,而这些方言与MSA的相似度各异,限制了模型迁移效果。本文通过3项自然语言处理任务和表征相似性分析,研究阿拉伯语模型的跨方言迁移。结果表明,迁移虽可行但存在显著差异,且部分可由方言地理邻近性解释。此外,训练支持所有阿拉伯方言的模型出现负向干扰现象,质疑方言间相似性的假设,引发对阿拉伯语模型跨语言迁移可靠性的担忧。
原文摘要 · Abstract (English)
Arabic Language Models (LMs) are pretrained predominately on Modern Standard Arabic (MSA) and are expected to transfer to its dialects. While MSA as the standard written variety is commonly used in formal settings, people speak and write online in various dialects that are spread across the Arab region. This poses limitations for Arabic LMs, since its dialects vary in their similarity to MSA. In this work we study cross-lingual transfer of Arabic models using probing on 3 Natural Language Processing (NLP) Tasks, and representational similarity. Our results indicate that transfer is possible but disproportionate across dialects, which we find to be partially explained by their geographic proximity. Furthermore, we find evidence for negative interference in models trained to support all Arabic dialects. This questions their degree of similarity, and raises concerns for cross-lingual transfer in Arabic models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。