用1.28亿乌尔都语令牌微调,让大模型更懂乌尔都语。
UrduLLaMA 1.0: Dataset Curation, Preprocessing, and Evaluation in Low-Resource Settings
- 基于Llama-3.1-8B架构,持续预训练1.28亿乌尔都语令牌。
- 在4.1万条指令和5万对翻译数据上用LoRA微调,性能超越现有模型。
- 适合研究低资源语言或需要多语言支持的开发者参考。
多语言大语言模型在乌尔都语等低资源语言上表现不佳。本文提出UrduLLaMA 1.0,基于开源Llama-3.1-8B-Instruct架构,在1.28亿乌尔都语令牌上进行持续预训练,捕捉语言多样性。为提升指令遵循与翻译能力,采用低秩适配(LoRA)在4.1万条乌尔都语指令和约5万对英乌翻译对上微调。在三个机器翻译数据集上的评估显示,其性能显著优于当前最先进(SOTA)模型,确立了乌尔都语大模型的新基准。研究结果表明,仅用有限数据与计算资源,通过针对性适配策略,即可有效应对低资源语言的独特挑战。
原文摘要 · Abstract (English)
Multilingual Large Language Models (LLMs) often provide suboptimal performance on low-resource languages like Urdu. This paper introduces UrduLLaMA 1.0, a model derived from the open-source Llama-3.1-8B-Instruct architecture and continually pre-trained on 128 million Urdu tokens, capturing the rich diversity of the language. To enhance instruction-following and translation capabilities, we leverage Low-Rank Adaptation (LoRA) to fine tune the model on 41,000 Urdu instructions and approximately 50,000 English-Urdu translation pairs. Evaluation across three machine translation datasets demonstrates significant performance improvements compared to state-of-the-art (SOTA) models, establishing a new benchmark for Urdu LLMs. These findings underscore the potential of targeted adaptation strategies with limited data and computational resources to address the unique challenges of low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。