构建72GB藏语语料库,持续预训练大模型提升藏语理解能力
From Curated Data to Scalable Models: Continual Pre-training of Dense and MoE Large Language Models for Tibetan
- 用72GB高质量藏语数据,持续预训练多语言大模型
- 密集与专家混合模型在藏语任务上均超越现有开源模型
- 为低资源语言建模提供可复用的方法,适合多语言研究者
大型语言模型在众多自然语言处理任务中表现卓越,但其性能严重偏向高资源语言。藏语虽具重要文化价值和庞大使用者群体,却仍严重缺失。本文提出一套完整流程,通过大规模数据整理与持续预训练推进藏语建模。我们构建了迄今最大的72GB高质量藏语语料库,采用平衡多语言持续预训练策略,对Qwen2.5-7B模型进行藏、汉、英三语训练,并进一步开展多语言指令微调。为更高效扩展容量,将密集模型扩展至50B-A10B的专家混合(MoE)架构。由于缺乏标准藏语评估基准,我们通过高质量翻译与人工验证构建多个评测数据集。实验表明,无论是密集模型还是MoE模型,在多样化任务中均持续优于同等规模的现有开源及藏语专用模型。本工作推动了以藏语为中心的大模型研究,并为其他低资源语言扩展提供了可迁移洞见。后续将公开模型权重、评测基准及详细数据处理文档。
原文摘要 · Abstract (English)
Large language models (LLMs) have achieved remarkable success across a wide range of natural language processing tasks, yet their performance remains heavily biased toward high-resource languages. Tibetan, despite its cultural significance and large speaker population, is still substantially underrepresented. In this work, we present a comprehensive pipeline for advancing Tibetan language modeling through large-scale data curation and continual pre-training. We construct a 72 GB high-quality Tibetan corpus, the largest to date, and adapt Qwen2.5-7B through balanced multilingual continual pre-training with Tibetan, Chinese, and English, followed by multilingual instruction tuning. To further scale capacity efficiently, we extend the dense model to a 50B-A10B Mixture-of-Experts architecture. Due to the absence of standardized Tibetan benchmarks, we build multiple evaluation datasets via high-quality translation and human verification. Experimental results show that both dense and MoE models consistently outperform existing open-source and Tibetan-focused models of similar scale across diverse tasks. Our work advances Tibetan-centric LLM research and provides transferable insights for extending LLMs to other low-resource languages. We will release the model weights, evaluation benchmarks, and detailed data processing documentation in the follow-up.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。