为马其顿语构建首个大规模语言模型与数据集,突破低资源语言瓶颈。
Towards Open Foundation Language Model and Corpus for Macedonian: A Low-Resource Language
- 构建3.5亿词的马其顿语语料库和10万条文化适配指令数据
- 训练80亿参数模型,在7项评测中超越所有同类模型,性能达10倍大模型水平
- 开源全部数据与模型,助力低资源语言研究与应用落地
全球技术普及带来对通用工具的需求,但大语言模型在低资源语言上能力受限。本文为马其顿语创建迄今最大语料库,含40GB文本、35亿词;构建包含10.6万条实例的指令数据集,注重文化适配性;建立涵盖七项基准的评估套件。基于此,训练出state-of-the-art的80亿参数模型domestic-yak,其在所有基准上优于现有8B级模型,表现接近10倍规模模型。母语者定性评估显示,该模型在语法准确性和文化契合度上更受青睐。所有数据、代码与模型权重已公开,可在github.com/LVSTCK获取源码,huggingface.co/LVSTCK获取预训练模型与数据。
原文摘要 · Abstract (English)
The increase in technological adoption worldwide comes with demands for novel tools to be used by the general population. Large Language Models (LLMs) provide a great opportunity in this respect, but their capabilities remain limited for low-resource languages, restricting applications in countries where such languages are spoken. We create several resources to facilitate the adoption of LLMs and to support research advancements for Macedonian. We collect the largest Macedonian corpus to date, consisting of 40GB of textual data and totaling 3.5B words. To support conversational applications, we collect a 106k-instance instruction dataset, carefully built to be culturally grounded. For evaluation, we construct a Macedonian evaluation suite covering seven benchmarks. Finally, we train domestic-yak, a state-of-the-art 8B-parameter model, on our curated datasets and evaluate it against eight baseline models using the newly constructed benchmark suite. Our model outperforms all existing models in the 8B parameter range across all benchmarks, and achieves performance comparable to models up to 10x larger. Furthermore, a qualitative analysis with native speakers reveals that our model is preferred over larger counterparts, receiving higher ratings for grammatical correctness and cultural appropriateness. All datasets, code, and model weights are openly released, setting a foundation for advancing LLMs in similarly underrepresented languages. These resources are publicly available at github.com/LVSTCK for source code, and at huggingface.co/LVSTCK for pretrained model weights and data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。