102M中文语料训练高效且符合认知规律的中文语言模型
The First ChineseBabyLM Challenge: training data-efficient and cognitively plausible language models for Chinese
- 用不超过102M中文词从零训练模型,不限架构与分词器
- 最优模型在理解、认知对齐和汉字知识三方面表现最佳
- 适合研究数据效率与认知合理性中文模型的研究者
本文介绍了作为NLPCC 2026一部分的首个中文婴儿语言模型挑战赛(ChineseBabyLM Challenge)。挑战要求参赛者使用不超过10200万中文词从头训练语言模型,评估涵盖自然语言理解、认知一致性与汉字知识三个维度。对分词器、模型架构及训练轮数无限制。共有18支队伍提交了28个不同模型,生成74份结果文件。优胜团队采用DeBERTa-v2架构,并在预训练中引入辅助拼音预测任务。部分提交还探索了课程学习策略与架构创新。该挑战为推进数据高效且认知合理的中文语言建模提供了基准。
原文摘要 · Abstract (English)
This paper presents the first ChineseBabyLM Challenge, organized as part of NLPCC 2026. The challenge asked participants to train language models from scratch using no more than 102M Chinese words. The models were evaluated on three tracks: natural language understanding, cognitive alignment, and Hanzi knowledge. There were no restrictions on tokenizers, model architectures, or the number of training epochs. Eighteen teams submitted 28 distinct models, generating 74 result files. The overall-winning team used a DeBERTa-v2 architecture and introduced an auxiliary pinyin-prediction objective during pretraining. Several submissions also explored curriculum-learning strategies and architectural innovations. Overall, the challenge provides a benchmark for advancing data-efficient and cognitively plausible approaches to Chinese language modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。