Sailor2是专为东南亚多语言打造的大型模型,支持13种语言且性能接近GPT-4o。
Sailor2: Sailing in South-East Asia with Inclusive Multilingual LLMs
- 基于Qwen2.5持续预训练5000亿词,其中4000亿为东南亚语料
- 200亿参数模型在东南亚语言上与GPT-4o打成平手(50-50胜率)
- 开源完整开发手册,适合想做小语种大模型的研究者
Sailor2 是一组面向东南亚(SEA)语言的先进多语言大模型,提供10亿、80亿和200亿三种规模以适配不同应用场景。该模型在Qwen2.5基础上,使用5000亿个令牌(其中4000亿为东南亚特定语料,1000亿为回放数据)进行持续预训练,支持13种东南亚语言,同时保持对中文和英文的熟练能力。Sailor2-20B模型在东南亚语言上与GPT-4o的对战胜率达到了50-50。我们还发布了一份全面的技术指南,涵盖数据整理、预训练、后训练、模型定制和评估五个关键环节。Sailor2模型(采用Apache 2.0许可证)旨在推动东南亚地区语言技术发展,其技术手册也期望激励研究者为其他低资源语言构建更具包容性的大模型。
原文摘要 · Abstract (English)
Sailor2 is a family of cutting-edge multilingual language models for South-East Asian (SEA) languages, available in 1B, 8B, and 20B sizes to suit diverse applications. Building on Qwen2.5, Sailor2 undergoes continuous pre-training on 500B tokens (400B SEA-specific and 100B replay tokens) to support 13 SEA languages while retaining proficiency in Chinese and English. Sailor2-20B model achieves a 50-50 win rate against GPT-4o across SEA languages. We also deliver a comprehensive cookbook on how to develop the multilingual model in an efficient manner, including five key aspects: data curation, pre-training, post-training, model customization and evaluation. We hope that Sailor2 model (Apache 2.0 license) will drive language development in the SEA region, and Sailor2 cookbook will inspire researchers to build more inclusive LLMs for other under-served languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。