罗马化输入让多语言模型表现更好,关键在共享子词单元。
Happiness is Sharing a Vocabulary: A Study of Transliteration Methods
- 用罗马化、音标等方法转换非拉丁文字,测试对模型影响。
- 罗马化在12项测试中11项胜出,显著优于其他方式。
- 共享更长的子词单元是提升模型性能的关键。
音译已成为弥合多语言自然语言处理中不同语言间差距的有前景方法,尤其适用于使用非拉丁字母的语言。本文研究了共享书写系统、重叠词元词汇和共享语音对多语言模型性能的影响。通过控制实验,采用三种音译方式(罗马化、音标转写、替换密码)及正字法作为输入,评估模型在命名实体识别(NER)、词性标注(POS)和自然语言推理(NLI)三项下游任务上的表现。结果表明,在12个评估设置中的11个,罗马化显著优于其他输入形式,与假设基本一致。进一步分析显示,与预训练语言共享更长的(子词)词元是有效利用模型能力的关键因素。
原文摘要 · Abstract (English)
Transliteration has emerged as a promising means to bridge the gap between various languages in multilingual NLP, showing promising results especially for languages using non-Latin scripts. We investigate the degree to which shared script, overlapping token vocabularies, and shared phonology contribute to performance of multilingual models. To this end, we conduct controlled experiments using three kinds of transliteration (romanization, phonemic transcription, and substitution ciphers) as well as orthography. We evaluate each model on three downstream tasks -- named entity recognition (NER), part-of-speech tagging (POS) and natural language inference (NLI) -- and find that romanization significantly outperforms other input types in 11 out of 12 evaluation settings, largely consistent with our hypothesis that it is the most effective approach. We further analyze how each factor contributed to the success, and suggest that having longer (subword) tokens shared with pre-trained languages leads to better utilization of the model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。