研究隐式推理过程如何随训练变化,发现高准确率下推理可能不忠实。
Final Checkpoints Are Not Enough: Analyzing Latent Reasoning Faithfulness Along Training Trajectories
- 通过反事实干预追踪训练中隐式推理的响应性变化。
- 准确率提升时,推理响应性反而下降,不同任务模式轨迹各异。
- 仅看最终模型会忽略推理忠诚度的关键演变,适合模型可解释性研究者。
隐式推理在连续隐藏状态中执行多步推断,有望实现更紧凑高效的推理。然而,这些黑箱状态引发忠实性问题:隐式推理步骤是否真正驱动最终答案?以往研究仅在特定检查点评估,发现多种不忠实行为。这种终点视角未能揭示忠实性如何随训练演变。本文通过在隐式推理状态上进行验证过的反事实编辑与干预,追踪行为和激活层面的证据。结果发现,高任务准确率可与低反事实响应性并存:随着准确率提升,响应性反而下降,不同隐式推理方法呈现不同演化轨迹。在ProsQA数据集上,输出对范数噪声替换的敏感性随反事实响应性下降,但结果依赖于替换方式。在独立训练的二选一与开放题GSM设置中,干预敏感性呈现相反趋势。这些结果表明,仅评估最终检查点会掩盖反事实响应性的变化时机及隐状态的真实贡献。
原文摘要 · Abstract (English)
Latent reasoning performs multi-step inference in continuous hidden states, promising more compact and efficient reasoning. However, these opaque states raise a question of faithfulness: whether the latent reasoning steps drive the final answer. Prior work studies this question at selected checkpoints and reports several unfaithful behaviors. This endpoint view leaves how evidence of faithfulness evolves during training unexamined. We track behavioral and activation-based evidence across training using verified counterfactual edits and interventions on the latent reasoning states. We find that high task accuracy can coexist with low counterfactual responsiveness: as accuracy improves, responsiveness can decline, and different latent reasoning approaches follow distinct trajectories. On ProsQA, output sensitivity to norm-noise replacement declines alongside counterfactual responsiveness, although the result depends on the replacement. Across separately trained binary-choice and open-ended GSM settings, intervention sensitivity follows opposite trajectories. These results show that evaluating only a final checkpoint can obscure both when counterfactual responsiveness changes and what the latent states contribute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。