用1万条指令数据提升哈萨克语大模型的政务文化理解能力
Instruction Tuning on Public Government and Cultural Data for Low-Resource Language: a Case Study in Kazakh
- 借助大模型生成10,600条高质量指令数据,覆盖哈萨克斯坦核心政务文化内容
- 在多项选择和生成任务中,微调Qwen/Falcon/Gemma均实现性能提升
- 适用于低资源语言的指令微调,尤其适合政府与文化领域应用
低资源语言的指令微调研究受限于文本数据不足,尤其在政务与文化领域。为此,我们构建并开源了一个大规模(10,600样本)的指令跟随(IFT)数据集,涵盖哈萨克斯坦关键机构与文化知识。该数据集提升了大模型对程序性、法律性及结构化治理内容的理解。采用大模型辅助生成,对比了开放权重与闭源模型的效果,最终选用GPT-4o作为生成骨干。每条数据均经人工全量验证以确保质量。实验表明,在本数据集上微调Qwen、Falcon与Gemma,可显著提升其在多项选择与生成任务中的表现,证明了大模型辅助指令微调在低资源语言中的潜力。
原文摘要 · Abstract (English)
Instruction tuning in low-resource languages remains underexplored due to limited text data, particularly in government and cultural domains. To address this, we introduce and open-source a large-scale (10,600 samples) instruction-following (IFT) dataset, covering key institutional and cultural knowledge relevant to Kazakhstan. Our dataset enhances LLMs' understanding of procedural, legal, and structural governance topics. We employ LLM-assisted data generation, comparing open-weight and closed-weight models for dataset construction, and select GPT-4o as the backbone. Each entity of our dataset undergoes full manual verification to ensure high quality. We also show that fine-tuning Qwen, Falcon, and Gemma on our dataset leads to consistent performance improvements in both multiple-choice and generative tasks, demonstrating the potential of LLM-assisted instruction tuning for low-resource languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。