用罗马化输入可显著提升多语言模型跨书写系统迁移能力。
One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography
- 在自回归预训练中使用罗马化输入,提升跨语言迁移效果。
- 罗马化预训练在8种语言上表现最佳,规模越大优势越明显。
- 罗马化应作为预训练核心设计,而非后期补救措施。
多语言语言模型通过共享子词词汇实现跨语言知识迁移,但当相关语言使用不同书写系统时,该机制失效。现有方法多采用字符等价(如罗马化或国际音标转写),但缺乏直接对比;且研究集中于编码器仅有的模型,多数工作是微调已有预训练模型。本文在控制条件下,系统比较了自回归多语言预训练中三种输入表示:正字法文本、国际音标(IPA)和罗马化,在三个规模(467M、709M、1.03B)下覆盖八种语言的四组类型学对。在多种下游任务(包括可见与不可见语言)上,罗马化预训练表现出最强跨语言迁移能力,且随规模扩大优势更明显。IPA在多数场景优于文本输入,但仍逊于罗马化。令人意外的是,将文本预训练模型在罗马化数据上微调会损害已覆盖语言的表现,仅在缺乏脚本覆盖时略有提升。结果表明,对于涵盖类型多样书写系统的多语言模型,为获得最大效益,应将罗马化作为预训练阶段的核心设计,而非后期修正。
原文摘要 · Abstract (English)
Multilingual language models transfer knowledge across languages through shared subword vocabulary, a mechanism that breaks down when related languages use different writing systems. Prior work addresses this via script equalization (romanization or IPA transcription), but direct comparisons are rare; the focus has been on encoder-only models, with most work adapting existing pretrained models. We systematically compare different input representations in autoregressive multilingual pretraining, comparing orthographic text, IPA, and romanization in a controlled setup across three scales (467M, 709M, and 1.03B) on eight languages in four typologically motivated pairs. Across a wide range of downstream tasks on seen and unseen languages, romanized pretraining yields the strongest cross-lingual transfer, and the advantage over text widens with scale. IPA improves over text in most settings but trails romanization. Surprisingly, finetuning a text-pretrained model on romanized data hurts performance on languages already covered by the base model, only marginally helping when the model lacks script coverage. Our results indicate that for multilingual models spanning typologically diverse scripts, to obtain maximum benefits, romanization should be treated as a core design choice applied at pretraining rather than a post hoc fix.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。