用数学指标量化大模型中意识般的全局激活现象。
The Ignition Index: Measuring Global Workspace Dynamics in Language Models

- 通过拟合探针准确率曲线,提取模型激活的突变程度。
- 不同架构的激活模式差异显著,如Mamba无全局广播特征。
- 适用于研究模型训练过程中的关键转折点,适合机制可解释性研究者。
我们提出了一种名为‘点火指数’(I)的验证性标量指标,用于在变压器语言模型中实现全局工作空间理论(GWT)的全或无点火预测。该指标对每层线性探针准确率随输入信号强度的变化进行四参数逻辑斯蒂拟合,提取陡度参数 β̂:高值表示突发式、类似点火的跃迁;低值则表示渐进式积累。在涵盖五种架构家族的11个模型中,打乱标签的对照实验表明,真实语言结构相对于虚假探针能力具有9.6倍的选择性(p < 0.001,Mann-Whitney U检验)。发现:(1) 前馈型变压器整体β̂值比状态空间模型(SSMs)高89%(p < 1e-13,Cohen's d = 0.52),Mamba表现出近似线性特征,与缺乏全局广播一致。(2) Huginn-3.5B在迭代轴上的点火程度是深度轴的2.12倍,说明循环架构在递归维度上呈现类工作空间跃迁。(3) Pythia-410M在训练步256处出现PELT检测到的相变(+67%),早于归纳头形成。(4) 模型规模和信号强度与点火的假设未被证实,提示变压器可能已饱和其可用的点火机制。点火指数首次建立了GWT动态预测与机械可解释性之间的定量桥梁,具备9.6倍测量选择性和此前未见的架构区分能力。
原文摘要 · Abstract (English)
We introduce the Ignition Index (I), a validated scalar metric that operationalizes Global Workspace Theory's (GWT) all-or-none ignition prediction in transformer language models. The metric fits a four-parameter sigmoid to per-layer linear probe accuracy as a function of input signal strength, extracting steepness parameter beta-hat: high values indicate abrupt, ignition-like transitions; low values indicate graded build-up. Across 11 models spanning five architecture families, shuffled-label controls demonstrate 9.6-fold selectivity for genuine linguistic structure over spurious probe capacity (p < 0.001, Mann-Whitney U-test). We find: (1) Feedforward transformers exceed SSMs by 89% in aggregate beta-hat (p < 1e-13, Cohen's d = 0.52), with Mamba exhibiting near-linear profiles consistent with absent global broadcast. (2) Huginn-3.5B exhibits 2.12-fold higher ignition along its iteration axis than its depth axis, demonstrating that recurrent architectures manifest workspace-like transitions along the recurrence dimension. (3) Pythia-410M shows a PELT-detected phase transition at training step 256 (+67%), preceding induction-head formation. (4) Hypotheses linking ignition to model scale and signal strength were not confirmed, suggesting transformer architectures may saturate available ignition mechanisms. The Ignition Index provides the first validated quantitative bridge between GWT's dynamical predictions and mechanistic interpretability, with 9.6-fold measurement selectivity and architecture-level discriminability not previously characterized in the scaling literature. Code: https://github.com/saman-rahbar/ignition-index
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。