打造首个覆盖11个东南亚语言的通用大模型,助力低资源语言AI发展。
SEA-LION: Southeast Asian Languages in One Network
- 通过多阶段微调与模型融合,构建支持11种东南亚语言的多语言大模型。
- 在多语言基准测试中表现优于现有同类模型,达到当前最佳水平。
- 开源模型资源,服务东南亚地区科研与应用需求。
大型语言模型(LLMs)在人工智能领域占据主导地位,但多数研究仍以英语为中心,导致东南亚(SEA)等低资源语言严重缺乏支持。为此,我们推出了Llama-SEA-LION-v3-8B-IT和Gemma-SEA-LION-v3-9B-IT两款先进多语言LLM,支持英语、中文、印尼语、越南语、马来语、泰语、缅甸语、老挝语、菲律宾语、泰米尔语和高棉语共11种东南亚语言。该工作采用大规模多语言持续预训练,并结合多阶段指令微调、对齐与模型合并策略。在多语言基准上的评估显示,我们的模型在支持东南亚语言的LLM中达到最先进性能。模型已开源,旨在惠及更广泛的东南亚社区。
原文摘要 · Abstract (English)
Recently, Large Language Models (LLMs) have dominated much of the artificial intelligence scene with their ability to process and generate natural languages. However, the majority of LLM research and development remains English-centric, leaving low-resource languages such as those in the Southeast Asian (SEA) region under-represented. To address this representation gap, we introduce Llama-SEA-LION-v3-8B-IT and Gemma-SEA-LION-v3-9B-IT, two cutting-edge multilingual LLMs designed for SEA languages. The SEA-LION family of LLMs supports 11 SEA languages, namely English, Chinese, Indonesian, Vietnamese, Malay, Thai, Burmese, Lao, Filipino, Tamil, and Khmer. Our work leverages large-scale multilingual continued pre-training with a comprehensive post-training regime involving multiple stages of instruction fine-tuning, alignment, and model merging. Evaluation results on multilingual benchmarks indicate that our models achieve state-of-the-art performance across LLMs supporting SEA languages. We open-source the models to benefit the wider SEA community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。