arXiv:2411.10083cs.CL2024-11

10亿参数多语言大模型,支持多种语言且性能领先。

Xmodel-1.5: An 1B-scale Multilingual LLM

  • 采用自研6.5万词表的unigram分词器,提升效率与准确率。
  • 在mMMLU、PIQA等评测中表现优异,泰语任务达顶尖水平。
  • 开源泰语专用数据集Xdata_Thai,助力低资源语言研究。

我们提出Xmodel-1.5,一个基于2万亿标记训练的10亿参数多语言大模型,兼顾性能与可扩展性。不同于多数使用BPE分词的模型,Xmodel-1.5采用自研的unigram分词器,拥有65,280个词元,优化了效率与准确性。该模型在泰语、阿拉伯语、法语、中文和英语等多种语言上表现均衡,其在对应评估数据集上的表现优于阿里云PolyLM-1.7B。在mMMLU和PIQA等基准测试中表现优异,并在泰语任务上达到当前最优水平。为支持低资源语言研究,我们发布了专用于泰语的Xdata_Thai数据集,涵盖性别词缀、习语等独特语言挑战。尽管模型表现良好,但在处理文化特异性语境方面仍有提升空间。模型与代码已公开于GitHub:https://github.com/XiaoduoAILab/XmodelLM-1.5。

原文摘要 · Abstract (English)

We introduce Xmodel-1.5, a 1-billion-parameter multilingual large language model pretrained on 2 trillion tokens, designed for balanced performance and scalability. Unlike most large models that use the BPE tokenizer, Xmodel-1.5 employs a custom unigram tokenizer with 65,280 tokens, optimizing both efficiency and accuracy. The model delivers competitive results across multiple languages, including Thai, Arabic, French, Chinese, and English, outperforming Alibaba's PolyLM-1.7B on respective evaluation datasets. Xmodel-1.5 excels in benchmarks like mMMLU and PIQA, and achieves state-of-the-art results in Thai. To support low-resource language research, we release Xdata_Thai, a Thai-specific evaluation dataset featuring unique linguistic challenges such as gendered particles and idioms. While the model demonstrates strong performance, there is still room for improvement in handling culturally specific nuances. We hope this work contributes to advancements in multilingual AI research. Models and code are publicly available on GitHub at https://github.com/XiaoduoAILab/XmodelLM-1.5

多语言模型分词器泰语开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。