实验证明,平行语料能显著提升大模型的多语言翻译与推理能力。
Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
- 通过控制实验验证平行数据对多语言能力的增强作用。
- 引入平行数据后,模型在翻译和多语言常识推理上表现明显提升。
- 适合关注多语言模型训练优化的研究者与工程师。
大型语言模型(LLMs)即使未显式训练于平行数据,也展现出出色的翻译能力。这促使部分研究者认为平行数据对构建多语言模型已不再必要。尽管有人归因于模型规模带来的涌现能力,但近期研究表明,实际原因是训练数据中存在偶然的双语信号。已有方法致力于最大化平行数据在多语言编码器及编码器-解码器模型中的效用。然而,部分解码器型LLM选择忽略平行数据。本文针对平行数据对LLMs多语言能力的影响进行系统性研究,重点关注翻译与多语言常识推理任务。通过受控实验,我们证明平行数据能显著提升LLMs的多语言能力。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated impressive translation capabilities even without being explicitly trained on parallel data. This remarkable property has led some to believe that parallel data is no longer necessary for building multilingual language models. While some attribute this to the emergent abilities of LLMs due to scale, recent work suggests that it is actually caused by incidental bilingual signals present in the training data. Various methods have been proposed to maximize the utility of parallel data to enhance the multilingual capabilities of multilingual encoder-based and encoder-decoder language models. However, some decoder-based LLMs opt to ignore parallel data instead. In this work, we conduct a systematic study on the impact of adding parallel data on LLMs' multilingual capabilities, focusing specifically on translation and multilingual common-sense reasoning. Through controlled experiments, we demonstrate that parallel data can significantly improve LLMs' multilingual capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。