arXiv:2606.01243cs.CLcs.LG2026-06ACL被引 2

通过可解释性指导,实现无需训练的推理过程干预

Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention

论文配图:Unlocking the Black Box of Latent Reasoning: An Interpretability-Guided Approach to Intervention
图 1 · 摘自论文原文
  • 用结构、因果、几何探测揭示隐式推理的内在机制
  • 在多个模型和任务上提升推理准确率,且不更新参数
  • 适合关注模型可解释性与可控推理的研究者

隐式推理使大语言模型在连续隐藏状态中完成多步推理,相比显式的思维链(CoT)更具效率。然而,这些连续思维向量的不可见性限制了其可靠性与可控性。本文通过系统性分析,结合结构、因果与几何探测,发现隐式向量编码了压缩且忠实的推理步骤信息,早期向量充当关键因果枢纽。基于此,我们设计了一套无需训练、解码时可执行的干预方法,通过施加识别出的几何与语义先验来优化隐式推理过程。在多种模型规模与任务领域上的实验证明,该方法能持续提升推理准确性。本研究实现了可解释性引导下的能力释放,显著改善推理表现,且无需任何参数更新。

原文摘要 · Abstract (English)

Latent reasoning enables Large Language Models (LLMs) to perform multi-step inference within continuous hidden states, offering efficiency gains over explicit Chain-of-Thought (CoT). However, the opacity of these continuous thought vectors hinders their reliability and controllability. This paper bridges the gap between mechanistic interpretability and actionable control. We first present a systematic analysis using structural, causal, and geometric probes, revealing that latent vectors encode compressed, faithful representations of reasoning steps, with early vectors acting as critical causal hubs. Building on this, we operationalize these interpretability insights into a suite of training-free, decode-time interventions that refine the latent reasoning process by imposing the identified geometric and semantic priors. Extensive experiments across multiple model scales and diverse task domains demonstrate that our approaches consistently improve reasoning accuracy. Our interpretability-guided interventions consistently unlock latent capabilities and improve reasoning accuracy without any parameter updates.

可解释性推理增强无训练干预

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。