arXiv:2412.13922cs.CLcs.AI2024-12被引 9

用6亿词语料持续预训练,让巴斯克语大模型指令理解能力提升24点。

Pipeline Analysis for Developing Instruct LLMs in Low-Resource Languages: A Case Study on Basque

  • 用高质量巴斯克语语料持续预训练,提升基础模型理解力
  • 自动翻译数据集实现指令微调与偏好对齐,性能提升24点
  • 首次在100亿参数以下建立巴斯克语新基准,适合低资源语言研究者

大型语言模型通常针对英语等资源丰富语言优化,加剧了高资源与低资源语言间的差距。本文以巴斯克语为例,详细分析了构建可执行指令的低资源语言模型的三个关键阶段:预训练、指令微调和人类偏好对齐。研究发现,使用约6亿词的高质量巴斯克语语料进行持续预训练,使基础模型的自然语言理解(NLU)得分提升超过12分。此外,利用自动翻译数据集进行指令微调和人类偏好对齐效果显著,使指令遵循性能提升24分。最终模型Llama-eus-8B和Llama-eus-8B-instruct在100亿参数以下类别中为巴斯克语建立了新的最先进水平。

原文摘要 · Abstract (English)

Large language models (LLMs) are typically optimized for resource-rich languages like English, exacerbating the gap between high-resource and underrepresented languages. This work presents a detailed analysis of strategies for developing a model capable of following instructions in a low-resource language, specifically Basque, by focusing on three key stages: pre-training, instruction tuning, and alignment with human preferences. Our findings demonstrate that continual pre-training with a high-quality Basque corpus of around 600 million words improves natural language understanding (NLU) of the foundational model by over 12 points. Moreover, instruction tuning and human preference alignment using automatically translated datasets proved highly effective, resulting in a 24-point improvement in instruction-following performance. The resulting models, Llama-eus-8B and Llama-eus-8B-instruct, establish a new state-of-the-art for Basque in the sub-10B parameter category.

低资源语言指令微调巴斯克语大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。