无需人工标注,用反向指令生成200种语言的高质量指令数据。
MURI: High-Quality Instruction Tuning Datasets for Low-Resource Languages via Reverse Instructions
- 通过反向指令和翻译管道从原文自动生成指令-输出对。
- 构建了超200万条指令数据,覆盖200种低资源语言。
- 适合研究多语言模型、资源匮乏语言应用的开发者使用。
指令微调能提升大语言模型在多种任务中与人类偏好的一致性。传统方法在低资源语言上受限于数据标注依赖。本文提出一种新方法——多语言反向指令(MURI),可在无需人工标注或预训练多语言模型的前提下,为低资源语言生成高质量指令微调数据集。该方法利用反向指令和翻译流程,从低资源语言的原始文本中生成指令-输出对,通过来自不同本土领域的文本来源和内容过滤机制,确保文化相关性和多样性。所构建的数据集MURI-IT包含超过200万条指令-输出对,覆盖200种语言。通过母语者评估及mT5模型的微调实验,验证了该方法在自然语言理解与开放式生成任务中的有效性。数据集与模型已公开发布于https://github.com/akoksal/muri。
原文摘要 · Abstract (English)
Instruction tuning enhances large language models (LLMs) by aligning them with human preferences across diverse tasks. Traditional approaches to create instruction tuning datasets face serious challenges for low-resource languages due to their dependence on data annotation. This work introduces a novel method, Multilingual Reverse Instructions (MURI), which generates high-quality instruction tuning datasets for low-resource languages without requiring human annotators or pre-existing multilingual models. Utilizing reverse instructions and a translation pipeline, MURI produces instruction-output pairs from existing human-written texts in low-resource languages. This method ensures cultural relevance and diversity by sourcing texts from different native domains and applying filters to eliminate inappropriate content. Our dataset, MURI-IT, includes more than 2 million instruction-output pairs across 200 languages. Evaluation by native speakers and fine-tuning experiments with mT5 models demonstrate the approach's effectiveness for both NLU and open-ended generation. We publicly release datasets and models at https://github.com/akoksal/muri.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。