arXiv:2412.05149cs.CL2024-12被引 70

用1亿词内数据训练语言模型,探索更接近人类学习效率的算法。

Findings of the Second BabyLM Challenge: Sample-Efficient Pretraining on Developmentally Plausible Corpora

  • 采用混合因果掩码架构提升小样本语言模型性能。
  • 训练计算量与多任务表现强相关,最佳方案改进了数据和目标函数。
  • 适合关注小数据语言建模、认知可解释性的研究者参考。

BabyLM挑战是社区推动缩小人类与计算机语言学习数据效率差距的尝试。参赛者需在不超过1亿词的数据预算下优化语言模型训练。今年发布了改进的文本语料库及视觉-语言语料库,以支持认知合理视觉语言模型研究。评估任务涵盖语法能力、(视觉)问答、语用能力与语义锚定等。参赛项目包括1000万词纯文本、1亿词纯文本和1亿词+图像多模态三类赛道。31个提交方案采用多样化方法,其中混合因果掩码语言模型架构表现最优。多模态赛道无方案超越基线。后续分析发现训练浮点运算量(FLOPs)与跨任务平均性能显著相关;表现最佳方案均对训练数据、目标函数和模型结构进行了调整。本届挑战表明该领域仍有巨大创新空间,尤其在图文建模方面,但社区协作仍能提供小规模语言建模的有效策略洞察。

原文摘要 · Abstract (English)

The BabyLM Challenge is a community effort to close the data-efficiency gap between human and computational language learners. Participants compete to optimize language model training on a fixed language data budget of 100 million words or less. This year, we released improved text corpora, as well as a vision-and-language corpus to facilitate research into cognitively plausible vision language models. Submissions were compared on evaluation tasks targeting grammatical ability, (visual) question answering, pragmatic abilities, and grounding, among other abilities. Participants could submit to a 10M-word text-only track, a 100M-word text-only track, and/or a 100M-word and image multimodal track. From 31 submissions employing diverse methods, a hybrid causal-masked language model architecture outperformed other approaches. No submissions outperformed the baselines in the multimodal track. In follow-up analyses, we found a strong relationship between training FLOPs and average performance across tasks, and that the best-performing submissions proposed changes to the training data, training objective, and model architecture. This year's BabyLM Challenge shows that there is still significant room for innovation in this setting, in particular for image-text modeling, but community-driven research can yield actionable insights about effective strategies for small-scale language modeling.

小样本学习语言模型多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。