发现大模型推理时有独特时空动态,可判断其思考是否有效。
Spatiotemporal Hidden-State Dynamics as a Signature of Internal Reasoning in Large Language Models

- 通过分析隐藏状态在解码步骤与层间的演变,捕捉推理过程
- 成功推理路径呈现广域时间动态与局部层集中特征
- 无需训练的StALT指标能准确区分正确与错误推理轨迹
大型推理模型(LRMs)生成长序列解题过程,但这些痕迹是否反映真实内部计算仍不明确。尽管近期研究显示隐藏状态包含正确性信号,但粗粒度聚合可能掩盖推理过程中的词元与层结构。本文研究解码步骤与层间隐藏状态的演化,发现成功推理路径具有广泛的时间动态和局部的层内集中特征,而非推理模型及知识密集型任务中该结构较弱。为此提出无需训练的时空转换幅度(StALT)指标,量化相邻词元间变化并加权于词元内的层显著性。在多种模型与基准测试中,StALT在推理密集场景下可靠地区分正确与错误轨迹,表现优于输出空间与长度基线。干预分析进一步表明,该指标对推理需求增减有系统响应,支持其与模型内部推理动力学的关联。结果为大模型存在可观测隐藏状态动态提供了实证,并提供了一种超越输出评估的内部计算探针。
原文摘要 · Abstract (English)
Large reasoning models (LRMs) generate extended solutions, yet it remains unclear whether these traces reflect substantive internal computation or merely verbosity and overthinking. Although recent hidden-state analyses suggest that internal representations carry correctness-related signals, their coarse aggregations may obscure the token and layer structure underlying reasoning computation. We investigate hidden-state transitions across decoding steps and layers, and identify a distinct spatiotemporal pattern in LRMs: successful trajectories exhibit broad temporal dynamics with localized layer-wise concentration, while this structure is weaker in non-reasoning models and knowledge-heavy domains. We formalize this characteristic as Spatiotemporal Amplitude of Latent Transition (StALT), a training-free trajectory statistic that summarizes temporal changes between adjacent tokens weighted by within-token layer saliency. Across diverse models and benchmarks, StALT reliably separates correct from incorrect trajectories in reasoning-intensive regimes, providing a competitive label-free correctness signal alongside strong output-space and length-based baselines. Intervention analyses further show that this spatiotemporal amplitude responds systematically to manipulations that increase or reduce the demand for internal reasoning, supporting its association with latent reasoning dynamics in LRMs. These findings provide empirical evidence that LRMs exhibit measurable hidden-state dynamics and offer a practical probe for understanding internal computation beyond output-based evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。