arXiv:2504.08165cs.CL2025-04被引 244

用少于1亿词数据训练出超越万亿词模型的语言模型

Findings of the BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

  • 在固定数据预算下优化训练策略,提升数据效率
  • 胜出模型仅用数亿词即超越万亿词训练的基准模型
  • 适合关注高效训练与认知建模的研究者

儿童仅需不到一亿词输入即可习得语言,而大语言模型通常需三到四个数量级更多数据,且表现仍不及人类。这种高资源需求限制了新模型训练及现有模型作为发展性认知模型的应用。BabyLM挑战赛是一项集体努力,参赛者在固定数据预算下竞争优化语言模型训练。评估涵盖语法能力、下游任务表现和泛化能力。参赛者可提交至三个数据限制逐步放宽的赛道。从30余份提交中,我们总结出提升数据效率的关键策略,并指出未来研究应重点关注的方向。采用LTG-BERT架构的优胜方案,在数亿词训练下表现超过使用万亿词训练的模型。其他方案通过短序列训练或学生-教师蒸馏也取得良好效果。尽管多数课程学习尝试未奏效,但少数略有提升。

原文摘要 · Abstract (English)

Children can acquire language from less than 100 million words of input. Large language models are far less data-efficient: they typically require 3 or 4 orders of magnitude more data and still do not perform as well as humans on many evaluations. These intensive resource demands limit the ability of researchers to train new models and use existing models as developmentally plausible cognitive models. The BabyLM Challenge is a communal effort in which participants compete to optimize language model training on a fixed data budget. Submissions are compared on various evaluation tasks targeting grammatical ability, downstream task performance, and generalization. Participants can submit to up to three tracks with progressively looser data restrictions. From over 30 submissions, we extract concrete recommendations on how best to train data-efficient language models, and on where future efforts should (and perhaps should not) focus. The winning submissions using the LTG-BERT architecture (Samuel et al., 2023) outperformed models trained on trillions of words. Other submissions achieved strong results through training on shorter input sequences or training a student model on a pretrained teacher. Curriculum learning attempts, which accounted for a large number of submissions, were largely unsuccessful, though some showed modest improvements.

数据效率语言模型认知建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。