为乌尔都语数学推理打造80亿参数大模型,填补低资源语言空白。
Riazi-8B: An Urdu Large Language Model for Mathematical Reasoning

- 在乌尔都维基百科上继续预训练,再用链式思维数据微调。
- 在MGSM-Urdu测试集上准确率显著优于现有乌尔都指令调优模型。
- 适合需要数学推理能力的乌尔都语用户或研究者使用。
近期的大语言模型在数学推理方面表现出色,但其进展主要依赖英文资源和基准测试。因此,在低资源语言如乌尔都语中,推理性能大幅下降,缺乏针对数学推理的数据集和适配模型。乌尔都语既无专门的推理资源,也缺少多步数学问题求解的模型,限制了最新成果对乌尔都语用户的适用性。我们通过两步适配过程构建了Riazi-8B:首先在乌尔都维基百科上进行持续预训练,然后在基于GSM8K生成的乌尔都链式思维数据上进行监督微调。我们在MGSM-Urdu上评估Riazi-8B,对比现有乌尔都指令调优模型,结果表明其在答案正确性、推理质量、响应完整性和乌尔都语生成方面均有稳定提升。研究证明,结合语言适配与推理导向微调是拓展低资源语言数学推理能力的有效策略。
原文摘要 · Abstract (English)
Recent LLMs demonstrate strong mathematical reasoning capabilities, but existing gains rely heavily on English-centric training resources and benchmarks. As a result, reasoning performance degrades substantially in low-resource languages such as Urdu, where reasoning-oriented datasets and adapted models remain scarce. Urdu lacks both reasoning-oriented resources and models adapted for multi-step mathematical problem solving, limiting the applicability of recent progress to Urdu-speaking users. We address this gap through Riazi-8B, an Urdu mathematical reasoning model developed through a two-step adaptation process comprising continued pre-training on Urdu Wikipedia and supervised fine-tuning on Urdu Chain-of-Thought data derived from GSM8K. We evaluate Riazi-8B on MGSM-Urdu against existing Urdu instruction-tuned models. Our results show consistent improvements in answer correctness, reasoning quality, response completeness, and Urdu generation. Our findings demonstrate that combining Urdu language adaptation with reasoning-focused fine-tuning is an effective strategy for extending mathematical reasoning capabilities to low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。