揭示大模型语法专精如何随训练逐步形成
How Syntax Specialization Emerges in Language Models
- 追踪模型训练全过程,分析语法敏感性演化轨迹
- 发现语法专精在特定层集中出现,存在关键发展期
- 结果跨架构与初始化稳定,适合模型机制研究者
大语言模型(LLMs)展现出令人意外的内部专精:个别神经元、注意力头和电路对句法结构表现出选择性敏感,反映出人类大脑中的模式。尽管这种专精已被广泛记录,其在训练过程中的形成机制及其影响因素仍不明确。本文通过追踪专精的演变过程,揭示其发展轨迹:语法敏感性逐步出现,在特定层集中显现,并经历快速内部专精的‘关键期’。该过程在不同架构和初始化参数(如随机种子)下保持一致,受模型规模和训练数据影响。因此,我们不仅揭示了语法在模型中的位置,还阐明了部分模型如何在训练中内化语法。为支持后续研究,论文将在接受后公开代码、模型及训练检查点。
原文摘要 · Abstract (English)
Large language models (LLMs) have been found to develop surprising internal specializations: Individual neurons, attention heads, and circuits become selectively sensitive to syntactic structure, reflecting patterns observed in the human brain. While this specialization is well-documented, how it emerges during training and what influences its development remains largely unknown. In this work, we tap into the black box of specialization by tracking its formation over time. By quantifying internal syntactic consistency across minimal pairs from various syntactic phenomena, we identify a clear developmental trajectory: Syntactic sensitivity emerges gradually, concentrates in specific layers, and exhibits a 'critical period' of rapid internal specialization. This process is consistent across architectures and initialization parameters (e.g., random seeds), and is influenced by model scale and training data. We therefore reveal not only where syntax arises in LLMs but also how some models internalize it during training. To support future research, we will release the code, models, and training checkpoints upon acceptance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。