平行语料对多语言表征对齐作用有限,早期训练中仅略提速。
On the limited utility of parallel data for learning shared multilingual representations
- 通过控制平行数据比例训练模型,研究其对跨语言对齐的影响。
- 无平行数据时跨语言表征对齐水平仍接近有数据情况,仅早期微调更快。
- 适合关注多语言预训练机制与数据效率的研究者阅读。
共享多语言表征对于跨语言任务和语言间知识迁移至关重要。本研究考察了平行数据(即翻译句子)在预训练中作为触发跨语言表征对齐信号的作用。我们训练了不同平行数据比例的基准模型,发现平行数据对跨语言对齐的影响极小。通过多种评估方法,我们发现其作用仅限于可能在预训练初期略微加速表示共享,并减少模型中的语言特异性神经元数量。即使没有平行数据的显式信号,跨语言对齐仍能以相似水平自然涌现。
原文摘要 · Abstract (English)
Shared multilingual representations are essential for cross-lingual tasks and knowledge transfer across languages. This study looks at the impact of parallel data, i.e. translated sentences, in pretraining as a signal to trigger representations that are aligned across languages. We train reference models with different proportions of parallel data and show that parallel data seem to have only a minimal effect on the cross-lingual alignment. Based on multiple evaluation methods, we find that the effect is limited to potentially accelerating the representation sharing in the early phases of pretraining, and to decreasing the amount of language-specific neurons in the model. Cross-lingual alignment seems to emerge on similar levels even without the explicit signal from parallel data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。