arXiv:2606.27460cs.CL2026-06

发现语言模型先学抽象统计规律,再学局部依赖关系。

Developmental approach reveals the statistical learning of Neural Language Models: Transformers generalize from the most abstract statistical patterns

论文配图:Developmental approach reveals the statistical learning of Neural Language Models: Transformers generalize from the most abstract statistical patterns
图 1 · 摘自论文原文
  • 用训练过程中的中间状态分析模型学习路径
  • 初期即掌握全局抽象统计,后期才补足局部依赖
  • 揭示模型认知与人类语言发展相似的渐进机制

本研究采用发展性方法探究神经语言模型(NLM)的统计学习与心理表征。一系列生成式Transformer模型在合成语法数据上训练,训练过程中保存多个阶段的模型状态。通过分析这些状态的内部表示变化,发现NLM在学习初期即获得最抽象的全局统计知识,随后逐步习得相对局部的统计依赖。整个学习过程包含大量早期过度泛化现象,这些泛化在后期逐渐被约束。基于此观察,我们提出一个新框架,用于解释NLM的统计学习与语言认知机制。

原文摘要 · Abstract (English)

In this study, we use a developmental approach to investigate the statistical learning and mental representation of neural language models (NLM). A series of Generative Transformer models are trained on a synthetic grammar. The model states are saved at multiple stages in the course of training. Through analyzing how the internal representations of these models change in the developmental path, we found that NLMs acquire the most abstract global statistical knowledge at the beginning of learning and later acquire the relatively local statistical dependencies. This learning path contains many over-generalizations from the very beginning and these over-generalizations are gradually constrained in the later stage of learning. Based on this observation, we propose a new framework to explain the statistical learning and language cognition of NLMs.

语言模型统计学习发展路径

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。