为2.3亿乌尔都语使用者打造的顶级大模型,性能超越前代44.64分
Qalb: Largest State-of-the-Art Urdu Large Language Model for 230M Speakers with Systematic Continued Pre-training
- 基于LLaMA-3.1 8B持续预训练19.7亿词元,融合英维数据防遗忘
- 在7项任务上达90.34分,较基线模型提升44.64分,超越此前最优
- 专为复杂乌尔都语设计,适合低资源语言研究与本地化应用
尽管大语言模型取得显著进展,但全球超过2.3亿使用者的乌尔都语在现代自然语言系统中仍严重缺失。现有多语言模型在乌尔都语任务上表现不佳,难以应对该语言复杂的形态学、从右到左的纳斯塔利克书写系统及丰富的文学传统。即使基础的LLaMA-3.1 8B-Instruct模型也难以生成流畅且符合语境的乌尔都语文本。我们提出Qalb,采用两阶段方法:持续预训练后进行监督微调。以LLaMA 3.1 8B为基础,在19.7亿词元的多样化乌尔都语数据集(涵盖新闻档案、古典与现代文学、政府文件及社交媒体)上进行持续预训练,并加入1.4亿词元的英文维基百科数据以防止灾难性遗忘。随后在Alif Urdu-instruct数据集上进行微调。在多项乌尔都语基准测试中,Qalb表现显著提升,加权平均得分达90.34,较此前最佳模型Alif-1.0-Instruct(87.1)高出3.24分,较基线模型提升44.64分。其在分类、情感分析、推理等七类任务上均达到当前最优水平。结果表明,通过高质量多样化语料的持续预训练结合针对性指令微调,可有效适配基础模型至低资源语言。
原文摘要 · Abstract (English)
Despite remarkable progress in large language models, Urdu-a language spoken by over 230 million people-remains critically underrepresented in modern NLP systems. Existing multilingual models demonstrate poor performance on Urdu-specific tasks, struggling with the language's complex morphology, right-to-left Nastaliq script, and rich literary traditions. Even the base LLaMA-3.1 8B-Instruct model shows limited capability in generating fluent, contextually appropriate Urdu text. We introduce Qalb, an Urdu language model developed through a two-stage approach: continued pre-training followed by supervised fine-tuning. Starting from LLaMA 3.1 8B, we perform continued pre-training on a dataset of 1.97 billion tokens. This corpus comprises 1.84 billion tokens of diverse Urdu text-spanning news archives, classical and contemporary literature, government documents, and social media-combined with 140 million tokens of English Wikipedia data to prevent catastrophic forgetting. We then fine-tune the resulting model on the Alif Urdu-instruct dataset. Through extensive evaluation on Urdu-specific benchmarks, Qalb demonstrates substantial improvements, achieving a weighted average score of 90.34 and outperforming the previous state-of-the-art Alif-1.0-Instruct model (87.1) by 3.24 points, while also surpassing the base LLaMA-3.1 8B-Instruct model by 44.64 points. Qalb achieves state-of-the-art performance with comprehensive evaluation across seven diverse tasks including Classification, Sentiment Analysis, and Reasoning. Our results demonstrate that continued pre-training on diverse, high-quality language data, combined with targeted instruction fine-tuning, effectively adapts foundation models to low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。