针对中文、印尼语等四语言优化的开源大模型,提升跨语言理解能力。
MERaLiON-TextLLM: Cross-Lingual Understanding of Large Language Models in Chinese, Indonesian, Malay, and Singlish
- 基于Llama-3-8B-base,通过持续预训练与权重融合优化
- 在四种语言基准测试中表现超越官方Llama-3模型
- 适合需要多语言支持的研究者与开发者使用
多语言大语言模型在多种语言中展现出强大能力,但在不同语系间表现差异显著,尤其对资源有限的语言。本文提出MERaLiON-TextLLM,一系列专为提升中文、印尼语、马来语和新加坡式英语理解与生成能力而设计的开源语言模型。初始发布模型基于Llama-3-8B-Base,通过精心设计的持续预训练与权重合并流程构建。该方法在相关语言基准测试中实现性能提升,超越官方Llama-3模型表现。我们公开模型权重,以支持跨语言语言理解的进一步研究与开发。
原文摘要 · Abstract (English)
Multilingual large language models (MLLMs) have shown impressive capabilities across a variety of languages. However, efficacy can differ greatly between different language families, especially for those with limited linguistic resources. This report presents MERaLiON-TextLLM, a series of open-source language models specifically tailored to improve understanding and generation in Chinese, Indonesian, Malay, and Singlish. The initial released model is built on Llama-3-8B-Base and refined through a meticulously crafted process of continued pre-training and weight merging. Our approach achieves performance improvements across benchmarks in these languages, exceeding the capabilities of the official Llama-3 models. We provide the model checkpoints as a resource to support further research and development in cross-lingual language understanding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。