行为表现不能准确反映模型推理能力发展,内部状态提前就已具备答案线索。
Behaviour Is an Incomplete Measure of Reasoning Development: Cross-surface pre-arrival accessibility and the limits of developmental inference in a recurrent-depth reasoner
- 通过隐藏状态探针发现模型在行为达标前已能预判答案
- 同一任务在不同训练表面上耗时相差186.5倍,但内部可访问性相似
- 研究揭示行为、内部状态与训练进展三者本质不同,适合认知科学与模型可解释性研究者
能力发展常通过行为阈值、最终检查点或解码器从隐状态中读出的信息推断。这些量未必对应同一事件。我们研究了一个30M参数的循环深度关系推理模型,在封闭且由预言定义的世界中,使用密集行为轨迹、两个训练表面、预先注册的到达前隐状态探针、前瞻性评估有效性验证及未训练和负向对照,全程区分训练时间轴与推理时间轴。行为层面:在固定获取标准下,符号表面三跳能力需70逻辑周期,而语言表面需13,055周期,差距达186.5倍;此后语言表面四跳能力在8逻辑周期内达成。在长达13,055周期的训练中,四跳保留行为始终未超过3/40,最终为0/40。内部测量显示:在语言表面,线性探针在行为到达前即恢复未来答案身份,得分0.056159(随机期望0.025),未训练对照0.024758,群体频率基线0.048309(p=0.012987;40个答案类别中有16个贡献)。类似提前可访问性在表面切换后仍存在,上游结构位置达0.1020(零步对照0.0460,p=0.000999),读出比较器处达0.0618(p=0.004),40类中有21类贡献。最后,自然尝试追踪该可访问性随训练变化不可行:探针有效性由行为到达定义,测量群体随被测对象变化。行为能力、内部可访问性与训练时间发展是三个独立可观测变量,行为与解码器可访问性均无法标识训练所获计算;因果干预是下一步必要步骤。
原文摘要 · Abstract (English)
Capability development is routinely inferred from behavioural thresholds, from final checkpoints, or from what a decoder can read out of a hidden state. These quantities need not identify the same event. We study a 30M-parameter recurrent-depth relational reasoner in a closed, oracle-defined world, using dense behavioural trajectories, two training surfaces, preregistered pre-arrival hidden-state probes, prospectively checked evaluability, and explicit untrained and negative controls, holding the training-time and inference-time axes separate throughout. Behaviour first: under one frozen acquisition criterion, three-hop competence cost 70 logical epochs on the symbolic surface and 13,055 on the verbal surface, a 186.5-fold contrast, after which verbal four-hop competence cleared in 8 logical epochs. Across the 13,055-epoch grind, four-hop held-out behaviour never exceeded 3/40 and ended at 0/40. Internal measurement next: on the verbal surface a linear probe recovered future-answer identity before behavioural arrival at 0.056159 against uniform chance 0.025, an untrained control of 0.024758 and a population frequency baseline of 0.048309 (p = 0.012987; 16/40 answer classes contributing). Analogous pre-arrival accessibility survived the surface change, reaching 0.1020 against a zero-step control of 0.0460 (p = 0.000999) at the upstream structural position and 0.0618 at the readout comparator (p = 0.004), with 21/40 classes contributing. Finally, the natural attempt to track that accessibility across training was not cleanly evaluable: probe eligibility is defined by behavioural arrival, so the measured population changes with the measurand. Behavioural competence, internal accessibility, and training-time development are distinct observables, and neither behaviour nor decoder accessibility identifies the computation training acquired; causal intervention is the necessary next step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。