arXiv:2605.15454cs.CLcs.LG2026-05

模型推理越长,路径越不同;修正长度影响后,难问题路径更直接。

Reasoning Models Don't Just Think Longer, They Move Differently

论文配图:Reasoning Models Don't Just Think Longer, They Move Differently
图 1 · 摘自论文原文
  • 用轨迹几何分析推理过程,剔除生成长度干扰。
  • 难度越高,修正后路径越直接,尤其在编程任务中明显。
  • 推理训练模型路径更稳定,适合研究推理机制的人看。

推理训练的语言模型在难题上往往生成更长的思维链,但更长的链并不说明模型在走不同的内部路径。我们通过竞赛编程、数学和布尔可满足性问题中的隐藏状态轨迹,研究了这一区别。原始轨迹几何受生成长度显著影响,不校正则难以比较难度差异。在校正长度后,所有领域中难度仍与修正后的轨迹几何系统相关。在代码领域,难题表现出更直接的修正轨迹和更低的局部曲率异质性,推理训练模型与指令微调基线存在明显差异。数学和布尔可满足性中关联较弱但仍存在。提示阶段的线性探测未复现代码领域的分离,行为标注显示更强的修正耦合伴随策略转变和不确定性监控。这些发现表明,长度校正是生成期轨迹分析的前提,并揭示推理训练可能带来特定的修正轨迹几何特征,其强度因任务域而异。

原文摘要 · Abstract (English)

Reasoning-trained language models often spend more tokens on harder problems, but longer chains of thought do not show whether a model is merely computing for more steps or following a different internal trajectory. We study this distinction through hidden-state trajectories during chain-of-thought generation across competitive programming, mathematics, and Boolean satisfiability. Raw trajectory geometry is strongly shaped by generation length: longer generations mechanically alter path statistics, so difficulty-dependent comparisons are misleading without adjustment. After residualizing trajectory statistics on length, difficulty remains systematically coupled to corrected trajectory geometry across all domains studied. The clearest reasoning-specific separation appears in the code domain, where harder problems show more direct corrected trajectories and less heterogeneous local curvature in reasoning-trained models than in matched instruction-tuned baselines. Corrected difficulty-geometry coupling is weaker, but still present, in mathematics and Boolean satisfiability. Prompt-stage linear probes do not mirror the code-domain separation, and behavioral annotations show that stronger corrected coupling co-occurs with strategy shifts and uncertainty monitoring. Together, these findings establish length correction as a prerequisite for generation-time trajectory analysis and show that reasoning training can be associated with distinct corrected trajectory geometry, with the strength of the effect depending on the domain.

推理模型轨迹分析思维链模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。