用少量芬兰语数据让大模型高效学会指令跟随。
Got Compute, but No Data: Lessons From Post-training a Finnish LLM
- 用英语数据翻译生成芬兰语指令,仅需几百样本即可训练
- 双语偏好优化使芬兰语表现最优,超越单语训练
- 适合低资源语言大模型适配与开源复现
随着大语言模型在聊天机器人和通用助手中的普及,指令跟随与人类偏好对齐的方法逐渐成熟,但其在高资源语言外的表现尚未验证。本文探讨了在英语和芬兰语上进行后训练的经验。我们使用多语言模型将英语指令和偏好数据翻译为芬兰语,分别在两种语言上进行指令微调和偏好优化,并评估模型在双语中的指令遵循能力。结果表明,仅需几百个芬兰语指令样本即可实现与主流模型相当的芬兰语指令遵循性能。此外,尽管英语偏好优化带来一定跨语言收益,但结合双语偏好数据能获得最佳效果。相关模型、数据集及训练方案已开源发布于 Hugging Face。
原文摘要 · Abstract (English)
As LLMs gain more popularity as chatbots and general assistants, methods have been developed to enable LLMs to follow instructions and align with human preferences. These methods have found success in the field, but their effectiveness has not been demonstrated outside of high-resource languages. In this work, we discuss our experiences in post-training an LLM for instruction-following for English and Finnish. We use a multilingual LLM to translate instruction and preference datasets from English to Finnish. We perform instruction tuning and preference optimization in English and Finnish and evaluate the instruction-following capabilities of the model in both languages. Our results show that with a few hundred Finnish instruction samples we can obtain competitive performance in Finnish instruction-following. We also found that although preference optimization in English offers some cross-lingual benefits, we obtain our best results by using preference data from both languages. We release our model, datasets, and recipes under open licenses at https://huggingface.co/LumiOpen/Poro-34B-chat-OpenAssistant
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。