用LoRA微调多语模型,发现中文生成变好但推理能力下降。
Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
- 用LoRA对Gemma模型微调马拉地语数据集
- 微调后评估指标下降但人工判断更优
- 适合研究低资源语言适配与评估方法
大型语言模型在多语言任务中表现卓越,但在低资源语言上的适配仍面临挑战。本研究探讨了低秩适应(LoRA)参数高效微调(PEFT)对多语Gemma模型在马拉地语(低资源语言)上的影响。基于包含52,000条指令-响应对的翻译版Alpaca数据集,实验发现尽管评估指标常显示微调后性能下降,但人工评估却普遍认为微调模型优于原始模型。结果表明,微调提升了目标语言生成能力,但削弱了模型的推理能力。该研究强调需改进评估方法,并构建高质量本地数据集,以更准确评估低资源语言下的模型表现。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable multilingual capabilities, yet challenges persist in adapting these models for low-resource languages. In this study, we investigate the effects of Low-Rank Adaptation (LoRA) Parameter-Efficient Fine-Tuning (PEFT) on multilingual Gemma models for Marathi, a language with limited resources. Using a translated Alpaca dataset with 52,000 instruction-response pairs, our findings reveal that while evaluation metrics often show a performance decline post-fine-tuning, manual assessments frequently suggest that the fine-tuned models outperform their original counterparts. The observations indicate improvements in target language generation capabilities but a reduction in reasoning abilities following language adaptation. These results underscore the need for improved evaluation methodologies and the creation of high-quality native datasets to accurately assess language-specific model performance in low-resource settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。