arXiv:2412.10244cs.CLcs.LG2024-12NAACL被引 12

低成本持续预训练,让大模型更好支持小语种

Efficient Continual Pre-training of LLMs for Low-resource Languages

  • 用新算法从语料库选精简文本,减少训练数据量
  • 仅用少量数据就能提升9种印度语言的生成效果
  • 适合资源匮乏语言的研究者和开发者使用

开源大语言模型虽赋予研究者灵活调整参数的能力,但在低资源语言(LRLs)上的表现仍逊于高资源语言(HRLs),主要因训练数据少、词汇覆盖不足。持续预训练(CPT)虽有效,但需大量特定语言数据,成本高昂。本文提出一种新算法,从大规模语料中筛选出最具代表性的文本子集,显著降低所需训练数据量;进一步设计新方法优化模型词汇表,仅通过有限增量即可提升性能。实验基于最新Llama-3模型,在9种印度语言(涵盖不同书写系统与资源水平)上进行,采用IndicGenBench作为生成任务评估基准。结果表明,使用小规模语料和扩展词汇表即可实现显著改进,且在不同语言家族间具可推广性。

原文摘要 · Abstract (English)

Open-source Large Language models (OsLLMs) propel the democratization of natural language research by giving the flexibility to augment or update model parameters for performance improvement. Nevertheless, like proprietary LLMs, Os-LLMs offer poorer performance on low-resource languages (LRLs) than high-resource languages (HRLs), owing to smaller amounts of training data and underrepresented vocabulary. On the other hand, continual pre-training (CPT) with large amounts of language-specific data is a costly proposition in terms of data acquisition and computational resources. Our goal is to drastically reduce CPT cost. To that end, we first develop a new algorithm to select a subset of texts from a larger corpus. We show the effectiveness of our technique using very little CPT data. In search of further improvement, we design a new algorithm to select tokens to include in the LLM vocabulary. We experiment with the recent Llama-3 model and nine Indian languages with diverse scripts and extent of resource availability. For evaluation, we use IndicGenBench, a generation task benchmark dataset for Indic languages. We experiment with various CPT corpora and augmented vocabulary size and offer insights across language families.

大模型低资源语言持续预训练印度语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。