arXiv:2412.13860cs.CLcs.LG2024-12被引 5

用合成数据微调大模型,让其适应低资源语言尼泊尔语。

Domain-adaptative Continual Learning for Low-resource Tasks: Evaluation on Nepali

  • 4-bit量化下用合成数据持续训练Llama 3,适配尼泊尔语
  • 评估显示模型在尼泊尔语生成上提升19.29%,基线仅4.98%
  • 通过注意力热图分析,发现模型保留了语法依赖解析能力

持续学习因大规模语言模型(LLMs)无法重新训练而成为重要研究方向。领域自适应预训练(DAPT)旨在持续训练预训练模型以适应未训练过的领域。本文评估了在低资源语言尼泊尔语上的DAPT可行性。采用合成数据,在4位量化QLoRA设置下继续训练Llama 3 8B模型以适配尼泊尔语。评估其性能、遗忘现象与知识获取能力。对比基础模型与最终模型在尼泊尔语生成、主流基准上的表现,并开展案例研究以探查其尼泊尔语语言知识。结果显示模型存在一定程度遗忘,但令人意外的是,增加评估时的提示数量可使最终模型性能提升高达19.29%,远超基线模型的4.98%,表明潜在知识保留。此外,通过层头自注意力热图分析,验证了最终模型在尼泊尔语中具备依赖关系解析能力。

原文摘要 · Abstract (English)

Continual learning has emerged as an important research direction due to the infeasibility of retraining large language models (LLMs) from scratch in the event of new data availability. Of great interest is the domain-adaptive pre-training (DAPT) paradigm, which focuses on continually training a pre-trained language model to adapt it to a domain it was not originally trained on. In this work, we evaluate the feasibility of DAPT in a low-resource setting, namely the Nepali language. We use synthetic data to continue training Llama 3 8B to adapt it to the Nepali language in a 4-bit QLoRA setting. We evaluate the adapted model on its performance, forgetting, and knowledge acquisition. We compare the base model and the final model on their Nepali generation abilities, their performance on popular benchmarks, and run case-studies to probe their linguistic knowledge in Nepali. We see some unsurprising forgetting in the final model, but also surprisingly find that increasing the number of shots during evaluation yields better percent increases in the final model (as high as 19.29% increase) compared to the base model (4.98%), suggesting latent retention. We also explore layer-head self-attention heatmaps to establish dependency resolution abilities of the final model in Nepali.

持续学习低资源语言大模型微调尼泊尔语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。