用少量数据和计算资源,让大模型学会加拿大法语方言。
Low-Resource Dialect Adaptation of Large Language Models: A French Dialect Case-Study
- 用低秩适配(LoRA)与高效持续预训练,仅更新1%参数。
- 在小数据集上适应魁北克法语,对主流法语任务影响极小。
- 成果可复现,开源首个魁北克法语LLM,惠及少数语言群体。
尽管大型语言模型(LLMs)已广泛使用,其强大能力仍主要局限于少数高资源语言。近期,持续预训练(CPT)被用于将模型微调至低资源地区方言。本文研究在严格数据与算力预算下,利用CPT进行方言学习的可行性。采用低秩适配(LoRA)与计算高效的持续预训练,我们仅用极小数据集,将三个LLM适配至魁北克法语,并在COLE基准上进行评测。实验表明,在仅更新约1%模型参数的情况下,少数方言任务性能显著提升,而主流语言任务性能几乎无损。结果分析显示,性能提升高度依赖语料构成。这些发现表明,结合参数高效微调(PEFT)的CPT能以低成本、可持续方式缩小方言差距,推动高质量大模型向少数语言社区普及。为支持可复现性与开放访问,我们已在Hugging Face发布首个魁北克法语大模型。
原文摘要 · Abstract (English)
Despite the widespread adoption of Large Language Models (LLMs), their strongest capabilities remain largely confined to a small number of high-resource languages for which there is abundant training data. Recently, continual pre-training (CPT) has emerged as a means to fine-tune these models to low-resource regional dialects. In this paper, we study the use of CPT for dialect learning under tight data and compute budgets. Using low-rank adaptation (LoRA) and compute-efficient continual pre-training, we adapt three LLMs to the Québec French dialect using a very small dataset and benchmark them on the COLE suite. Our experiments demonstrate an improvement on the minority dialect benchmarks with minimal regression on the prestige language benchmarks with around 1% of model parameters updated. Analysis of the results demonstrate that gains are highly contingent on corpus composition. These findings indicate that CPT with parameter-efficient fine-tuning (PEFT) can narrow the dialect gap by providing cost-effective and sustainable language resource creation, expanding high-quality LLM access to minority linguistic communities. To support reproducibility and broaden access, we release the first Québec French LLMs on Hugging Face.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。