用罗马化文本预训练多语言模型,对拼音文字影响小,对汉字日文有损失但可缓解。
One Script Instead of Hundreds? On Pretraining Romanized Encoder Language Models
- 在六种高资源语言上对比罗马化与原文预训练,评估信息损失与跨语言干扰。
- 拼音文字性能几乎无损,汉字日文性能下降,高保真罗马化可减轻但无法完全恢复。
- 罗马化提升拼音文字编码效率,适合追求高效多语言建模的场景。
揭示潜在词法重叠后,罗马化成为提升多语言模型跨语言迁移能力的有效策略。以往研究多集中于高资源拉丁语系向低资源非拉丁语系迁移,或亲缘关系近的语言间转换,因此尚不清楚罗马化是否适用于通用多语言模型预训练,尤其是其带来的信息损失是否会影响高资源语言表现。本文在六种语言类型多样、高资源的语言上从头预训练编码器语言模型,分别使用罗马化文本和原始文本,考察两类潜在退化因素:(i) 脚本特异性信息丢失,(ii) 因词汇重叠增加引发的负向跨语言干扰。采用两种不同保真度的罗马化工具,发现拼音文字性能几乎无损,而音节文字(如中文、日文)出现性能下降,更高保真的罗马化虽能缓解但无法完全恢复。进一步对比单语言模型与多语言模型,未发现子词重叠导致负面干扰。此外,罗马化显著提升拼音文字的编码效率(即‘肥力’),且性能代价可忽略。
原文摘要 · Abstract (English)
Exposing latent lexical overlap, script romanization has emerged as an effective strategy for improving cross-lingual transfer (XLT) in multilingual language models (mLMs). Most prior work, however, focused on setups that favor romanization the most: (1) transfer from high-resource Latin-script to low-resource non-Latin-script languages and/or (2) between genealogically closely related languages with different scripts. It thus remains unclear whether romanization is a good representation choice for pretraining general-purpose mLMs, or, more precisely, if information loss associated with romanization harms performance for high-resource languages. We address this gap by pretraining encoder LMs from scratch on both romanized and original texts for six typologically diverse high-resource languages, investigating two potential sources of degradation: (i) loss of script-specific information and (ii) negative cross-lingual interference from increased vocabulary overlap. Using two romanizers with different fidelity profiles, we observe negligible performance loss for languages with segmental scripts, whereas languages with morphosyllabic scripts (Chinese and Japanese) suffer degradation that higher-fidelity romanization mitigates but cannot fully recover. Importantly, comparing monolingual LMs with their mLM counterpart, we find no evidence that increased subword overlap induces negative interference. We further show that romanization improves encoding efficiency (i.e., fertility) for segmental scripts at a negligible performance cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。