arXiv:2606.26050cs.LGcond-mat.dis-nn2026-06

模型学了规则又突然忘记,谁输谁赢取决于训练数据中规则出现频率。

Natural Ungrokking: Asymmetric Control of Which Rules Survive Pretraining

  • 规则能否留存由训练数据中该规则出现频次决定
  • 规则消失时概率优势在100步内归零,且不可逆
  • 大模型更易遗忘,小模型可被精准控制删除规则

在一次常规预训练过程中,小型语言模型在第925步学会代词-性别规则(如'Sue cried because'后接'she'),在未见测试中得分达0.94。但到第3500步时,同一模型在相同测试上得分接近零,尽管训练数据仍包含该规则的证据。这种训练过程中的规则消退称为自然未领悟(natural ungrokking):模型的保留规则由语料自行决定,损失曲线无迹可寻。规则的存亡可由一个语料统计量预测——规则在训练流中获胜的频率。在多个语料、预算和随机种子下,支持频率决定规则命运;数据量与参数量之比仅调节被遗忘程度。相同先显现后崩溃的现象也出现在公开的Pythia检查点中,遗忘深度随模型规模递增,符合预测。遗忘是替代过程:一种竞争性表面模式取代原规则,两者对数概率差在行为崩溃前100步内跨过零点。控制权不对称:破坏规则的操作无法恢复它。在原处反转支持为反证,能单调剂量响应地消除两个独立规则;但即使注入支持至自然维持水平的450倍,也无法恢复规则。所有验证阈值与预测均在读取数据前注册。

原文摘要 · Abstract (English)

Midway through an ordinary pretraining run, a small language model learns the pronoun-gender rule: cued with a girl's name ("Sue cried because"), it resolves the next pronoun to she, generalizing to held-out probes (0.94 by step 925). By step 3,500 the same model scores near zero on the same probes, although the rule's evidence is still in the training data. We call this within-run reversal natural ungrokking: the corpus decides, with no trace in the loss curve, which learned rules a model keeps. Which rules survive is predictable from one corpus statistic: how often the training stream shows the rule winning. Across un-intervened runs (two corpora, three budgets, three seeds), support frequency decides a rule's fate; the data-to-parameter ratio only modulates how deeply a doomed rule falls. The same emerge-then-collapse dynamics appear in public Pythia checkpoints, collapse depth ordered by model scale as predicted. The forgetting is a displacement: a competing surface pattern out-competes the rule, and the log-probability margin between them crosses zero within 100 training steps of the behavioral collapse. Control over this fate is asymmetric: the same edit that destroys a rule on demand cannot restore it. Flipping support to counter-evidence in place kills the rule with monotone dose-response in two unrelated rules; but injecting support back, even to 450 times the level that naturally sustains it, buys no recovery. Every confirmatory threshold and prediction was pre-registered before the data it governed was read.

模型遗忘规则学习预训练机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。