揭示大模型跨语言迁移失败的根源,提出统一表征训练方法。
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
- 通过合成多语言数据训练小模型,研究事实与语言的关联性
- 仅当跨语言表征统一时,知识才能有效迁移
- 为提升大模型跨语言能力提供可操作的训练策略
大型语言模型在跨语言知识迁移中表现不佳:当以一种语言提问、而事实在另一语言中训练时,模型可能产生幻觉。本文通过从零训练小型Transformer模型,在可控的合成多语言数据集上研究该现象的原因与训练动态。根据(1)事实与其所用语言的相关性(信息量),及(2)语言识别的难易程度(可提取性),模型会形成统一或分离的跨语言表征;只有在表征统一时,知识才能实现跨语言迁移。基于这些发现,本文提出统一视角,解释了多语言大模型中一系列先前观察到的现象。研究表明,可控实验能揭示预训练动态,并建议将促进表征统一的方法纳入训练流程,从而改善大模型的跨语言迁移能力。
原文摘要 · Abstract (English)
Large language models (LLMs) struggle with cross-lingual knowledge transfer: they sometimes hallucinate when asked in one language about facts expressed in a different language during training. This work introduces a controlled setting to study the causes and training dynamics of this phenomenon by training small Transformer models from scratch on synthetic multilingual datasets. Depending on (1) the correlation between facts and the language they were learned in (informativeness), and (2) the ease of language identification (extractability), models either develop unified representations across languages or separate representations; only when representations are unified do facts transfer across languages. Based on these insights, we propose a unifying perspective which explains a range of prior observations concerning cross-lingual transfer in multilingual LLMs. Our work shows controlled settings can shed light on pre-training dynamics and suggests methods to encourage representational unification as part of training that would improve LLMs' cross-lingual transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。