arXiv:2510.09051cs.CLcs.AI2025-10中稿 · EMNLP被引 5

用合成数据训练出性能优越的乌尔都语大模型,成本低于100美元。

Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation

  • 通过改进自指导方法生成高质量多语言合成数据
  • 在不到100美元预算内超越多个主流大模型
  • 适合低资源语言研究者与本土化AI开发

为低资源语言如乌尔都语开发高性能大语言模型面临数据稀缺、多语言不一致和安全问题等挑战。现有方法依赖大量翻译数据,但质量差且成本高。本文提出Alif-1.0-8B-Instruct,基于预训练Llama-3.1-8B,采用改进的自指导技术构建高质量多语言合成数据集(Urdu-Instruct)。该数据集通过独特提示与种子值设计,结合全局任务池,融入乌尔都语原生思维链推理、双语翻译、文化相关性及伦理安全对齐。模型在乌尔都语特定任务上表现优于Llama-3.1-8B-Instruct,同时超越Mistral-7B-Instruct-v0.3、Qwen-2.5-7B-Instruct和Cohere-Aya-Expanse-8B,在训练成本低于100美元的前提下实现卓越性能。所有数据、模型与代码已公开。

原文摘要 · Abstract (English)

Developing a high-performing large language models (LLMs) for low-resource languages such as Urdu, present several challenges. These challenges include the scarcity of high-quality datasets, multilingual inconsistencies, and safety concerns. Existing multilingual LLMs often address these issues by translating large volumes of available data. However, such translations often lack quality and cultural nuance while also incurring significant costs for data curation and training. To address these issues, we propose Alif-1.0-8B-Instruct, a multilingual Urdu-English model, that tackles these challenges with a unique approach. We train the model on a high-quality, multilingual synthetic dataset (Urdu-Instruct), developed using a modified self-instruct technique. By using unique prompts and seed values for each task along with a global task pool, this dataset incorporates Urdu-native chain-of-thought based reasoning, bilingual translation, cultural relevance, and ethical safety alignments. This technique significantly enhances the comprehension of Alif-1.0-8B-Instruct model for Urdu-specific tasks. As a result, Alif-1.0-8B-Instruct, built upon the pretrained Llama-3.1-8B, demonstrates superior performance compared to Llama-3.1-8B-Instruct for Urdu specific-tasks. It also outperformed leading multilingual LLMs, including Mistral-7B-Instruct-v0.3, Qwen-2.5-7B-Instruct, and Cohere-Aya-Expanse-8B, all within a training budget of under $100. Our results demonstrate that high-performance and low-resource language LLMs can be developed efficiently and culturally aligned using our modified self-instruct approach. All datasets, models, and code are publicly available at: https://github.com/traversaal-ai/alif-urdu-llm.

乌尔都语合成数据多语言低成本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。