arXiv:2606.12689cs.CL2026-06被引 2

发现可观察的潜在模式不等于内部推理机制,需用因果测试验证。

Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models

论文配图:Observable Patterns Are Not Explanations: A Causal-Geometric Analysis of Latent Reasoning Models
图 1 · 摘自论文原文
  • 通过对比实验,发现控制组也出现类似推理模式
  • 潜在思考对行为的影响是渐进的,非二元开关
  • 适合关注模型可解释性与因果分析的研究者

潜变量推理模型(LRMs)以连续思维替代显式思维链。近期研究将可观测的潜在状态模式(如广度优先搜索类前沿、可解码的算术运算)视为内部推理机制的证据。我们对两种LRM(Coconut和CODI)与缺乏所提循环结构或课程训练的控制组进行对比,发现这些模式同样出现在控制组中,且并不总对行为产生因果影响。因果干预表明,潜在思维的利用程度并非二值化,而是随其对模型行为的因果效应呈梯度变化。几何分析显示,这种效应集中于低秩方向,且随着行为影响增强,其步骤间几何结构更加有序。因此,潜变量思维应被视为隐藏计算,而非隐藏解释:可解码性、注意力分布或静态结构本身无法确立机制。LRM可解释性需依赖匹配控制组与因果检验。

原文摘要 · Abstract (English)

Latent reasoning models (LRMs) replace explicit chain-of-thought with continuous thoughts. Recent work treats observable latent-state patterns, such as BFS-like frontiers and decodable arithmetic computation, as evidence for internal reasoning mechanisms. Evaluating two LRMs (Coconut and CODI) against controls lacking the proposed recurrence or curriculum, we find these patterns also appear in the controls and do not always causally affect behavior. Causal interventions reveal that latent-thought utilization is not binary but graded, scaling with a thought's causal effect on model behavior. Geometric analyses reveal this effect concentrates in low-rank directions whose step-to-step geometry grows more structured as their behavioral influence increases. Latent thoughts should therefore be treated as hidden computation, not hidden explanation: decodability, attention, or static structure alone cannot establish mechanism. LRM interpretability thus requires matched controls and causal tests.

模型可解释性因果分析潜变量模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。