AlignX通过双阶段对齐提升多语言大模型表现
AlignX: Advancing Multilingual Large Language Models with Multilingual Representation Alignment
- 第一阶段用语义对齐与语言特征融合对齐多语言表征
- 第二阶段通过多语言指令微调激发模型跨语言能力
- 显著改善非主流语言性能,适合多语言应用开发者
多语言大语言模型具备出色的多语言理解和生成能力,但在非主导语言上的表现和跨语言对齐仍不理想。常见方法是使用大规模、更均衡的多语言语料进行微调,但往往导致对齐不精确且知识迁移效果不佳,各语言提升有限。本文提出AlignX,一种两阶段的表征级框架,用于提升预训练多语言大模型的多语言性能。第一阶段通过多语言语义对齐与语言特征融合对齐多语言表征;第二阶段通过多语言指令微调增强模型的多语言能力。在多个预训练模型上的实验表明,该方法显著提升了模型的多语言通用性与跨语言生成能力。进一步分析显示,AlignX使多语言表征更接近,改善了跨语言对齐。
原文摘要 · Abstract (English)
Multilingual large language models (LLMs) possess impressive multilingual understanding and generation capabilities. However, their performance and cross-lingual alignment often lag for non-dominant languages. A common solution is to fine-tune LLMs on large-scale and more balanced multilingual corpus, but such approaches often lead to imprecise alignment and suboptimal knowledge transfer, struggling with limited improvements across languages. In this paper, we propose AlignX to bridge the multilingual performance gap, which is a two-stage representation-level framework for enhancing multilingual performance of pre-trained LLMs. In the first stage, we align multilingual representations with multilingual semantic alignment and language feature integration. In the second stage, we stimulate the multilingual capability of LLMs via multilingual instruction fine-tuning. Experimental results on several pre-trained LLMs demonstrate that our approach enhances LLMs' multilingual general and cross-lingual generation capability. Further analysis indicates that AlignX brings the multilingual representations closer and improves the cross-lingual alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。