arXiv:2512.02213cs.LG2025-12

为非洲低资源语言构建高质量指令数据集,提升大模型对话能力。

InstructLR: A Scalable Approach to Create Instruction Dataset for Under-Resourced Languages

  • 用大模型生成+双层过滤,保证语料质量
  • 产出三个各含5万条指令的多领域数据集
  • 适合低资源语言研究者和本地化应用开发者

当前大型语言模型在低资源语言(LRLs)上的文本生成与对话系统仍面临挑战,主要源于高质量指令数据集的匮乏,这一问题在非洲地区及其他区域的语言中尤为突出。现有方法如自动翻译和合成数据生成,常导致输出缺乏流畅性或拼写一致性。本文提出InstructLR框架,通过结合大模型生成与双层质量过滤机制——基于检索增强生成(RAG)的n-shot提示自动化筛选,以及人工验证层——有效提升数据质量。借鉴MMLU等基准的任务定义思路,InstructLR已成功构建三个多领域指令数据集:ZarmaInstruct-50k、BambaraInstruct-50k和FulfuldeInstruct-50k,每套数据均包含5万条高质量指令。

原文摘要 · Abstract (English)

Effective text generation and chat interfaces for low-resource languages (LRLs) remain a challenge for state-of-the-art large language models (LLMs) to support. This is mainly due to the difficulty of curating high-quality instruction datasets for LRLs, a limitation prevalent in the languages spoken across the African continent and other regions. Current approaches, such as automated translation and synthetic data generation, frequently yield outputs that lack fluency or even orthographic consistency. In this paper, we introduce InstructLR, a novel framework designed to generate high-quality instruction datasets for LRLs. Our approach integrates LLM-driven text generation with a dual-layer quality filtering mechanism: an automated filtering layer based on retrieval-augmented-generation (RAG)-based n-shot prompting, and a human-in-the-loop validation layer. Drawing inspiration from benchmarks such as MMLU in task definition, InstructLR has facilitated the creation of three multi-domain instruction benchmarks: ZarmaInstruct-50k, BambaraInstruct-50k, and FulfuldeInstruct-50k.

低资源语言指令数据集大模型非洲语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。