arXiv:2606.29164cs.LGcs.AI2026-06

发现语言模型推理轨迹中存在稳定不变的方向,可提升推理一致性。

Invariant Reasoning Directions in Latent Trajectories of Language Models

论文配图:Invariant Reasoning Directions in Latent Trajectories of Language Models
图 1 · 摘自论文原文
  • 通过对比强弱推理轨迹,识别出低秩的稳定推理方向。
  • 干预这些方向使推理一致性提升约10%,轨迹方差减少50%。
  • 无需训练,适合研究模型内在推理机制或改进推理稳定性。

隐式推理模型在隐藏状态空间中直接进行多步推理,但其隐式推理轨迹的结构仍不明确。我们发现,强弱推理轨迹间的对比信号具有高度集中的低秩结构,而无约束的隐状态更新则对改写、检查点选择和轨迹扰动敏感。这表明隐式推理轨迹包含稳定的不变方向与不稳定的实例相关变化。为此,我们提出无需训练的干预框架TILR,首先从跨输入的对比轨迹差异中学习一个低秩不变子空间,再通过自适应对齐门控抑制不匹配更新。在六个推理基准上,少量隐式方向解释了强弱推理轨迹间的主要差异。对这些方向的干预能因果性地提升推理一致性,并减少改写和扰动下的轨迹不稳定性。TILR在改写下使答案一致性提升约10%,轨迹方差降低至多50%,同时保持推理准确率。结果支持一种几何视角:可迁移的推理行为源于隐藏状态轨迹中的稳定低维结构。

原文摘要 · Abstract (English)

Latent reasoning models perform multi-step inference directly in hidden-state space, yet the structure of these latent reasoning trajectories remains poorly understood. We show that contrastive refinement signals between stronger and weaker reasoning trajectories exhibit a highly concentrated low-rank structure, while unconstrained latent updates remain sensitive to paraphrases, checkpoint choice, and trajectory perturbations. These observations suggest that latent reasoning trajectories contain stable invariant directions mixed with unstable instance-specific variation. We introduce \textbf{Trajectory-Invariant Latent Refinement (TILR)}, a training-free intervention framework for identifying and manipulating stable reasoning directions in latent space. TILR first learns a low-rank invariant subspace from contrastive trajectory differences across inputs, then constrains latent interventions to this subspace while suppressing poorly aligned updates through an adaptive alignment gate. Across six reasoning benchmarks, we find that a small number of latent directions explain most variation between strong and weak reasoning trajectories. Interventions on these directions causally improve reasoning consistency and reduce trajectory instability under paraphrases and perturbations. TILR improves answer consistency under paraphrase by ~10% and reduces latent trajectory variance by up to $50\%$ while preserving reasoning accuracy. These results support a geometric view of latent reasoning in which transferable reasoning behavior emerges from stable low-dimensional structure within hidden-state trajectories.

语言模型推理轨迹不变性隐空间干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。