arXiv:2601.17664cs.CLcs.AI2026-01

为乌尔都语构建轻量级大模型,性能媲美30倍大的多语言模型

UrduLM: A Resource-Efficient Monolingual Urdu Language Model

  • 用33GB自建语料训练1亿参数单语模型,降低计算开销
  • 少样本下情感分类准确率达66.6%,语法纠错BLEU超30
  • 开源全部资源,助力低资源语言研究

乌尔都语全球使用者达2.3亿,但缺乏专用的Transformer语言模型和高质量语料库。现有多语言模型对乌尔都语支持有限,存在性能差、计算成本高及文化偏差等问题。为此,我们提出UrduLM,一个在低资源环境下预训练的乌尔都语单语语言模型。通过整合多样化来源,构建了33GB乌尔都语语料库;开发定制化BPE分词器,相较多语言方案减少至少20%-30%的分词开销;并预训练了一个1亿参数的解码器模型。在少样本评估中,UrduLM性能可媲美规模高达其30倍的多语言模型,在情感分类任务上达到66.6%准确率,语法纠错任务的BLEU分数超过30。完整方法——包括语料、分词器、模型权重和评估基准——均已公开,旨在为乌尔都语自然语言处理研究建立基线,并为其他低资源语言提供可扩展框架。

原文摘要 · Abstract (English)

Urdu, spoken by 230 million people worldwide, lacks dedicated transformer-based language models and curated corpora. While multilingual models provide limited Urdu support, they suffer from poor performance, high computational costs, and cultural inaccuracies due to insufficient training data. To address these challenges, we present UrduLM, a pretrained Urdu monolingual language model trained in low-resource settings. We curate a 33GB Urdu corpus from diverse sources, develop a custom BPE tokenizer that reduces tokenization overhead by atleast 20-30% compared to multilingual alternatives, and pretrain a 100M-parameter decoder-only model. In few-shot evaluations, UrduLM achieves competitive performance with multilingual models up to 30x its size, reaching 66.6% accuracy on sentiment classification and BLEU scores exceeding 30 on grammar correction tasks. The complete methodology -- including corpus, tokenizer, model weights, and evaluation benchmarks -- is released openly to establish a baseline for Urdu NLP research and provide a scalable framework for other underrepresented languages.

语言模型低资源乌尔都语开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。