arXiv:2501.03855cs.CL2025-01

用少量数据训练的BabyLM模型在科萨语上表现优异,验证了低资源语言可行性。

BabyLMs for isiXhosa: Data-Efficient Language Modelling in a Low-Resource Context

  • 在少于1亿词的科萨语文本上预训练两种高效模型
  • 命名实体识别F1提升3.2点,部分超越XLM-R
  • 适合数据稀缺但需高效建模的低资源语言研究

BabyLM挑战要求开发样本高效的语言模型,参赛模型仅在小于1亿词的英语语料上预训练(相当于儿童语言发展暴露量)。该挑战催生了新型数据高效建模范式,其性能优于在数万亿词上训练的模型。本文探索此类模型在低资源语言中的潜力,以科萨语为例。我们在科萨语文本上预训练了两种BabyLM架构:ELC-BERT和MLSM,结果在词性标注和命名实体识别任务上均优于基础预训练模型,其中后者在命名实体识别上取得+3.2 F1的显著提升。在某些情况下,其表现甚至超过XLM-R。研究证实数据高效模型在低资源语言中具备可行性,但也凸显高质量预训练数据的匮乏。最后,我们对BabyLM架构如何编码科萨语进行了可视化分析。

原文摘要 · Abstract (English)

The BabyLM challenge called on participants to develop sample-efficient language models. Submissions were pretrained on a fixed English corpus, limited to the amount of words children are exposed to in development (<100m). The challenge produced new architectures for data-efficient language modelling, which outperformed models trained on trillions of words. This is promising for low-resource languages, where available corpora are limited to much less than 100m words. In this paper, we explore the potential of BabyLMs for low-resource languages, using the isiXhosa language as a case study. We pretrain two BabyLM architectures, ELC-BERT and MLSM, on an isiXhosa corpus. They outperform a vanilla pretrained model on POS tagging and NER, achieving notable gains (+3.2 F1) for the latter. In some instances, the BabyLMs even outperform XLM-R. Our findings show that data-efficient models are viable for low-resource languages, but highlight the continued importance, and lack of, high-quality pretraining data. Finally, we visually analyse how BabyLM architectures encode isiXhosa.

低资源语言高效建模预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。