让低资源语言更好用,通过跨语言表征对齐提升模型表现
ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework
- 将非主导语言表征移入主导语言空间,再移回原语言
- 在低资源语言上性能显著提升,最高提升17.3个点
- 适合多语言模型优化与低资源语言应用研究
尽管使用多语言数据微调大语言模型可快速提升其多语言能力,但因各语言训练数据不均衡,仍存在主导语言(如英语)与非主导语言间的性能差距。为进一步提升非主导语言表现,我们提出ShifCon——一种基于迁移的多语言对比框架,通过将非主导语言的内部表示映射至主导语言子空间,使其能利用模型参数中丰富的信息。随后,将增强后的表示移回原始语言子空间以生成内容。我们引入子空间距离度量,定位最优迁移层,并采用多语言对比学习强化该区域内的表示对齐。实验表明,该框架显著提升了非主导语言性能,尤其在低资源语言上效果突出。深入分析验证了方法有效性,为未来研究提供新思路。
原文摘要 · Abstract (English)
Although fine-tuning Large Language Models (LLMs) with multilingual data can rapidly enhance the multilingual capabilities of LLMs, they still exhibit a performance gap between the dominant language (e.g., English) and non-dominant ones due to the imbalance of training data across languages. To further enhance the performance of non-dominant languages, we propose ShifCon, a Shift-based multilingual Contrastive framework that aligns the internal forward process of other languages toward that of the dominant one. Specifically, it shifts the representations of non-dominant languages into the dominant language subspace, allowing them to access relatively rich information encoded in the model parameters. The enriched representations are then shifted back into their original language subspace before generation. Moreover, we introduce a subspace distance metric to pinpoint the optimal layer area for shifting representations and employ multilingual contrastive learning to further enhance the alignment of representations within this area. Experiments demonstrate that our ShifCon framework significantly enhances the performance of non-dominant languages, particularly for low-resource ones. Further analysis offers extra insights to verify the effectiveness of ShifCon and propel future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。