小数据训练语言模型,支持多语种,推动认知建模与AI结合
BabyLM Turns 4 and Goes Multilingual: Call for Papers for the 2026 BabyLM Workshop
- 用百万级以下语料训练模型,分大/小规模两赛道
- 新增英、荷、中文多语种训练,拓展跨语言研究
- 适合关注高效训练与认知科学融合的研究者
BabyLM旨在激发认知建模与语言模型预训练之间的新研究连接。我们邀请相关成果投稿至BabyLM研讨会,本届还将举办第四届BabyLM挑战赛。延续往年设置,挑战赛包含两个标准赛道:严格(Strict)和严格小规模(Strict-Small),分别要求在1亿词以下和1000万词以下的语料上训练语言模型。今年首次突破纯英语数据限制,新增多语种赛道,聚焦英语、荷兰语和中文。研讨会征稿主题涵盖训练效率、小规模数据集、认知建模、模型评估及架构创新等方向。
原文摘要 · Abstract (English)
The goal of the BabyLM is to stimulate new research connections between cognitive modeling and language model pretraining. We invite contributions in this vein to the BabyLM Workshop, which will also include the 4th iteration of the BabyLM Challenge. As in previous years, the challenge features two ``standard'' tracks (Strict and Strict-Small), in which participants must train language models on under 100M or 10M words of data, respectively. This year, we move beyond our previous English-only pretraining datasets with a new Multilingual track, focusing on English, Dutch, and Chinese. For the workshop, we call for papers related to the overall theme of BabyLM, which includes training efficiency, small-scale training datasets, cognitive modeling, model evaluation, and architecture innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。