arXiv:2608.17120cs.CL2026-08

孩子学词越来越快,而语言模型始终稳定增速。

Children, but not language models, show accelerating returns in word learning

  • 孩子词汇增长呈加速积累,每份新语言经验带来的收获递增。
  • 语言模型学习速率恒定,符合数据规模的线性增长规律。
  • 揭示儿童高效学习机制,适合教育科技与认知研究者参考。

儿童在生命最初几年学习数百个词汇,这一过程起初缓慢,随后迅速加速。以往模型将词汇增长视为随时间累积证据的过程。本文表明,该过程更应被描述为加速累积:每个新增的语言经验所带来词汇收获,都超过前一次。相比之下,语言模型——即使以儿童语言数据训练——并未表现出加速,而是呈现对新数据的恒定比例收益,符合缩放定律。儿童使用的训练数据量比语言模型少多个数量级;其对学习输入效率的不断提升,或可解释这一差异。

原文摘要 · Abstract (English)

Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.

儿童语言学习效率语言模型认知科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。