混合单语预训练模型合并会引发性能崩溃,因表示相似性不足。
On the Limits of Model Merging for Multilinguality in Pre-Training
- 测试单语模型合并的可行性,发现存在严重干扰
- 合并后多语言性能全面下降,远低于联合预训练
- 强调表示相似性是模型合并的前提,适合研究多语言融合
赋予模型一致的多语言性能可通过混合预训练数据或语言特定的模型合并实现。本文检验了将合并方法应用于单语预训练模型的可行性。在控制实验中,对比了混合、合并与单语预训练三种设置。结果表明,虽然单语预训练在本语言上表现优异,但任意组合的单语模型合并均导致性能崩溃,源于显著的干扰。分析显示,表示相似性是模型合并的必要前提。因此,合并方法在微调中的灵活性无法直接扩展到语言特定的预训练阶段。
原文摘要 · Abstract (English)
Endowing models with consistent multilingual performance can be achieved by mixing pre-training data, or post-training approaches such as language-specific model merging. In this work, we test whether merging can be applied to monolingually pre-trained models. We conduct a controlled study on the efficacy of mixed, merged, and monolingual pre-training setups. We find that while monolingual pre-training results in strong in-language performance, merging any combination of monolingual models leads to performance collapse due to interference. Our analysis suggests representational similarity is a prerequisite for model merging. We therefore conclude that the flexibility of merging in fine-tuning does not extend trivially to language-specific pre-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。