arXiv:2606.02378cs.LGcs.AI2026-06

揭示10亿参数模型中注意力电路形成的时间规律与机制差异。

When Do Attention Circuits Form? Developmental Trajectories of Capability and Attention-Sink Emergence Across Three 1B-ClassArchitectures

  • 通过参与度比率和选择性筛选,追踪三类模型的注意力头演化轨迹。
  • 诱导电路形成早于注意力陷阱形成,相差10-20倍训练词数。
  • 不同架构下注意力陷阱出现方式各异,有的渐进、有的突变。

我们跟踪了三种10亿参数级语言模型(Pythia 1B、OLMo 1B-0724-hf、OLMoE 1B-7B-0924)在两个架构家族(密集Transformer、专家混合)和两个预训练语料库(The Pile、DCLM)下的注意力头电路演化过程。每模型在10个对数间隔的版本上进行分析,共30次可解释性实验。采用参与度比率(PR)谱信号与全头能力特异性选择性筛选,监测诱导、前词依赖及BOS吸引子头的形成。发现:(F1) 所有模型中第0、1层均无BOS分类头,此为架构特性。(F2) 整体模型的BOS吸引子比例呈现三种不同形态:Pythia 1B为渐进上升,OLMo 1B从7%跃升至70%(相邻检查点),OLMoE 1B-7B为渐进上升。(F3) 在DCLM语料训练的模型中,诱导电路形成比BOS吸引子早10-20倍训练词数;能力电路与注意力陷阱是两个独立演化阶段。(F4) 能力特异性筛选在总训练词数的0.3%-2%内收敛至最终诱导电路,无需等待最终模型。(F5) 每个最终诱导头在首次超过能力选择性阈值时,其参与度比率即显著升高。结果表明,在10亿参数模型中,诱导相变与注意力陷阱相变在训练量上相差一个数量级,且形态迥异。

原文摘要 · Abstract (English)

We track the developmental trajectory of attention-head circuit formation across three 1B-class language models spanning two architecture families (dense transformer, mixture-of-experts) and two pretraining corpora (The Pile, DCLM): Pythia 1B, OLMo 1B-0724-hf, and OLMoE 1B-7B-0924. At each of 10 log-spaced revisions per model -- 30 mechanistic-interpretability runs in total -- we apply a participation-ratio (PR) spectral signal and an all-head capability-specific selectivity screen to track induction, previous-token, and BOS-attractor heads as they emerge. Five findings. (F1) Layers 0 and 1 produce zero BOS-classified heads at every revision in every model: the L0/L1 zero-BOS floor is an architectural property, not a learned outcome. (F2) The whole-model BOS-attractor fraction follows three distinct emergence shapes -- a gradual ramp in Pythia 1B, a sharp phase transition in OLMo 1B (7% to 70% between adjacent checkpoints), and a gradual ramp in OLMoE 1B-7B. (F3) In DCLM models, induction-circuit formation precedes BOS-attractor formation by 10-20x in tokens; capability-circuit formation and attention-sink formation are two transitions, not one. (F4) The capability-specific screen converges to the final induction circuit within 0.3-2% of total training tokens -- circuit identification does not require the final model. (F5) For every final-checkpoint induction head sampled across all three models, per-head PR is elevated at or before the first revision at which that head crosses its capability-selectivity threshold. The results refine the induction-phase-transition framing: in 1B-class models trained on DCLM, the induction transition and the attention-sink transition are separated by an order of magnitude in tokens and have qualitatively different shapes.

注意力机制模型演化可解释性语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。