用平行语料低成本构建首个真正开源的东南亚语言大模型
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
- 仅用平行语料进行持续预训练,提升多语言能力
- 34.7B token数据+180小时8卡H200,性能媲美同类模型
- 首次公开训练数据,适合研究模型透明性与偏见
大型语言模型在自然语言处理任务中表现优异,但多数仍以英语为中心,对低资源语言效果不佳。尽管已有面向东南亚的语言模型,但均非真正开源,未公开训练数据。真正开源模型对透明度及深入理解模型内部机制(如偏见、泛化、多语言性)至关重要。受近期研究表明平行语料可有效提升多语言性能的启发,我们开展控制性实验,验证其在持续预训练中的作用。结果表明,仅使用平行语料是扩展模型至新语言最有效的方法。仅用34.7亿词元平行数据和8张NVIDIA H200 GPU上180小时训练,我们构建了OpenSeal——首个真正开源的东南亚语言大模型,性能达到同规模模型水平。
原文摘要 · Abstract (English)
Large language models (LLMs) have proven to be effective tools for a wide range of natural language processing (NLP) applications. Although many LLMs are multilingual, most remain English-centric and perform poorly on low-resource languages. Recently, several Southeast Asia-focused LLMs have been developed, but none are truly open source, as they do not publicly disclose their training data. Truly open-source models are important for transparency and for enabling a deeper and more precise understanding of LLM internals and development, including biases, generalization, and multilinguality. Motivated by recent advances demonstrating the effectiveness of parallel data in improving multilingual performance, we conduct controlled and comprehensive experiments to study the effectiveness of parallel data in continual pretraining of LLMs. Our findings show that using only parallel data is the most effective way to extend an LLM to new languages. Using just 34.7B tokens of parallel data and 180 hours on 8x NVIDIA H200 GPUs, we built OpenSeal, the first truly open Southeast Asian LLM that rivals the performance of existing models of similar size.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。