arXiv:2412.13375cs.CL2024-12中稿 · COLING 2025被引 7

让大模型学会波斯语,用高效微调提升多语言能力

Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation

  • 分三阶段微调:波斯语单语预训练、双语对齐、任务专用指令微调
  • 波斯语分类准确率提升,英语任务性能无损甚至略有提高
  • 数据少时模型初始能力更重要,跨语言迁移对复杂任务帮助有限

大语言模型在分类和文本生成任务上取得显著进展,但主要基于英语数据训练,对低资源语言表现不佳。本研究以波斯语为例,探索通过参数高效微调将新语言融入Llama模型(原对波斯语理解有限)。采用多阶段方法:首先在单语波斯语数据上预训练,再通过双语预训练和指令数据对齐表示,最后使用特定任务数据进行指令微调。在每个阶段评估生成与分类任务性能。结果表明,通过双语数据对齐加入波斯语可提升波斯语分类准确率,且对英语任务无负面影响,有时还带来改善。此外,初始模型能力是有限数据下的关键因素,跨语言对齐对低资源语言帮助有限;英语到波斯语的知识迁移影响微弱,主要惠及简单分类任务。

原文摘要 · Abstract (English)

Large language models (LLMs) have made great progress in classification and text generation tasks. However, they are mainly trained on English data and often struggle with low-resource languages. In this study, we explore adding a new language, i.e., Persian, to Llama (a model with a limited understanding of Persian) using parameter-efficient fine-tuning. We employ a multi-stage approach involving pretraining on monolingual Persian data, aligning representations through bilingual pretraining and instruction datasets, and instruction-tuning with task-specific datasets. We evaluate the model's performance at each stage on generation and classification tasks. Our findings suggest that incorporating the Persian language, through bilingual data alignment, can enhance classification accuracy for Persian tasks, with no adverse impact and sometimes even improvements on English tasks. Additionally, the results highlight the model's initial strength as a critical factor when working with limited training data, with cross-lingual alignment offering minimal benefits for the low-resource language. Knowledge transfer from English to Persian has a marginal effect, primarily benefiting simple classification tasks.

多语言微调波斯语Llama

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。