语言模型能力涌现像一次罕见事件,全电路对齐才触发,部分正确无用。
The Kinetics of Training: A Driven-Nucleation Rate Law for Emergence, Plasticity Loss, and Circuit Control in Language Models

- 能力出现需所有电路部分一次性对齐,缺一不可。
- 删掉一个关键模块仅保留17%能力,远低于预期的50%-83%。
- 重置查询-键参数可恢复学习能力,证明机制可干预。
语言模型中某一能力的出现依赖于其电路最后部分在一次随机尝试中全部对齐,部分正确毫无意义。我们证明这种‘零部分分’的联合对齐是能力形成的速率限制步骤。在无捷径装置中,五部件电路缺三与三部件缺三等待时间相当(1.19-1.37),说明等待时间取决于缺失数量而非整体大小;在Pythia模型上,七种能力、三种规模下,移除一个组件后,32个判别单元中能力中位值仅剩17%(32/32显著低于预测的50-83%,p=2e-10),而随机头则保留100%。单一罕见事件且其障碍随缺失部件增加而上升,遵循速率方程:站点×尝试×驱动×exp(-βK)减去破坏项,可正向预测能力点燃时刻,反向揭示学习延迟但验证损失持续下降,传统监控失效。定位损伤源头为头部对基础数据的锁定,并提出修复方案:仅重置查询-键切片可恢复学习(6/6成功),价值切片无效(0/6)。在受控门控注意力模型中验证机制:占据强制设定截止时间,无需混合假设。最终通过引入符合涨落-耗散定理的噪声并渐进调节,成功熔断、固定电路。研究范围涵盖变压器中的合取电路至1.4B参数规模。
原文摘要 · Abstract (English)
A capability appears in a language model when the last parts of its circuit align in one stochastic attempt, and getting all but one right is worth nothing. We show this no-partial-credit joint alignment is the rate-limiting step of capability formation. Two fingerprints: in a shortcut-free apparatus a five-part circuit missing three waits as long as a three-part circuit missing three (1.19-1.37), so the wait counts missing parts, not size; and on Pythia across seven capabilities and three scales, ablating one part leaves a median 17% of the capability in 32 of 32 discriminating cells, where partial credit predicts 50-83% (p = 2e-10), while a random non-part head leaves 100%. One rare event whose barrier grows with missing parts yields a rate equation -- sites x attempts x drive x exp(-beta*K), minus destruction -- read three ways, each preregistered with frozen constants. Forward: a capability flat at baseline ignites at a step of our choosing once the mix passes a concentration floor (10/10 above, 0/12 below), and while still flat its arrival is datable from its precursor to 5% median error on six held-out models. Backward: the delay to learn a withheld capability grows with waiting until, past a critical step, it never ignites -- yet validation loss falls smoothly throughout, so standard monitors are blind to it. We locate the damage (heads commit to the base data) and isolate the cure: re-initializing only the query-key slices restores learnability (6/6) while the value slices do nothing (0/6). We prove the mechanism in a controlled gated-attention model: occupation forces a deadline whose consequences need no mixing assumption. Completed: SGD's noise fails the fluctuation-dissipation test, so we install one and anneal, melt and pin circuits on schedule. Scope: conjunction circuits in transformers to 1.4B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。