arXiv:2608.03263cs.LGcs.AI2026-08被引 1

发现推理模型的组合点燃现象真实存在于输出层,且与问题难度严格对应。

The Ignition Is Real, and It Lives at the Readout: Latent composition, difficulty-clocked ignition, and the interface-constituted commit in a recurrent-depth reasoner

  • 通过独立复现模型并追踪训练过程,验证了组合点燃的真实性。
  • 决策置信度在提交时刻突增5.8-8.0个逻辑值,远超阈值附近非事件步的96%。
  • 隐藏状态方向迅速锁定,后续变化以径向为主,输出层无显著影响。

我们检验了隐空间推理模型中报告的“组合点燃”是否为真实计算、仪器误差或源于语言训练数据。我们从头开始复现一个3000万参数的递归深度推理模型(相同配方与种子),全程记录其发展过程,通过预注册的全签名门验证一致性,并同时测量词汇输出层和隐藏状态的解析能力。结果显示,点燃现象真实存在且位于输出层:到达时间随问题深度规律上升,解析能力清晰且稳定,签名在两个同种子但不同训练轨迹的实现中重现。在提交时刻,决策置信度在一次迭代中跃升5.8-8.0个逻辑值,超过90百分位近阈值非事件步的96%;此时符号边际的零交叉仅为定义性,不具证据权重,因此证据来自条件幅度。隐藏状态方向在原始几何中迅速锁定,满足预注册标准(在解码器的LayerNorm坐标下衰减略低于阈值,故复合解码器坐标声明未被确认),随后在描述上冻结(解码器坐标中角步从52.9降至1.2度,历时八次迭代),后续位移主要为径向(平方范数占比0.961),输出层影响极小(径向逻辑值效应≤5.7e-6)。早期速度谷点声称被撤回:预注册归一化控制显示其依赖坐标。中间表示无法通过耦合输出层恢复(中继值0.00)。所有标准均在数据前冻结。预测日志,包括本文自身撤回的标题,已发布于配套仓库。

原文摘要 · Abstract (English)

We test whether the "compositional ignition" reported in latent-reasoning models is real computation, an instrument artifact, or inherited from verbal training data. We grow an independent realization of a published 30M-parameter recurrent-depth reasoner from scratch (same recipe and seed), film its development, certify fidelity through a pre-registered whole-signature gate, and measure resolution in two channels at once: the vocabulary readout and the hidden state. The ignition is real and lives at the readout: arrival time rises lawfully with problem depth, resolution is sharp and holds, and the signature reproduces across two same-seed realizations with divergent training trajectories. At commitment the decision margin jumps 5.8-8.0 logits in one iteration, exceeding the 90th percentile of near-threshold non-event steps in 96% of cases; the signed margin's zero-crossing there is definitional and carries no evidential weight, so the evidence is that conditioned magnitude. The hidden-state direction snaps in raw geometry, meeting its pre-registered criterion (in the decoder's LayerNorm coordinates it attenuates just below our bar, so the composite decoder-coordinate claim is not confirmed), and then freezes in both (descriptively so in decoder coordinates; angular steps 52.9 to 1.2 degrees over eight iterations), while subsequent displacement is predominantly radial (0.961 of squared-norm) and readout-null to a measured bound (radial logit effect <=5.7e-6). An earlier velocity-trough claim is withdrawn: pre-registered normalization controls showed it coordinate-dependent. Intermediates were never recoverable through the tied readout (relay 0.00). All criteria were frozen before their data; the predictions ledger, including this paper's own withdrawn headline, ships in the companion repository.

深度推理组合点燃模型机制可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。