LLM推理本质是潜在状态演变,而非表面思维链。
LLM Reasoning Is Latent, Not the Chain of Thought
- 将推理视为潜在状态轨迹演化,而非显式思维链。
- 实验证明潜在状态变化比表面文本更关键,提升推理性能。
- 适合研究模型内在机制或设计更精准评估方法的学者。
本文主张,大语言模型(LLM)推理应被理解为潜在状态轨迹的形成,而非忠实的表面思维链(CoT)。这一区分至关重要,因为关于可解释性、推理基准和推理时干预的有效性等讨论,均依赖于对推理核心对象的定义。本文提出三个可竞争的假设:H1(推理主要由潜在状态轨迹驱动)、H2(推理主要由显式表面思维链驱动)和H0(大部分推理提升可归因于通用串行计算,而非特定表征)。在分离常被混淆的三类因素基础上,重新组织近期实证、机制与调查研究,并引入经计算审计的示例,分解表面痕迹、潜在干预与匹配预算扩展,结果表明当前证据最支持以H1为默认工作假设。因此,本文建议:第一,将潜在状态动态作为研究推理的默认对象;第二,设计能明确解耦表面痕迹、潜在状态与串行计算的评估方案。
原文摘要 · Abstract (English)
This position paper argues that large language model (LLM) reasoning should be studied as latent-state trajectory formation rather than as faithful surface chain-of-thought (CoT). This matters because claims about faithfulness, interpretability, reasoning benchmarks, and inference-time intervention all depend on what the field takes the primary object of reasoning to be. We ask what that object should be once three often-confounded factors are separated and formalize three competing hypotheses: H1, reasoning is primarily mediated by latent-state trajectories; H2, reasoning is primarily mediated by explicit surface CoT; and H0, most apparent reasoning gains are better explained by generic serial compute than by any privileged representational object. Reorganizing recent empirical, mechanistic, and survey work under this framework, and adding compute-audited worked exemplars that factorize surface traces, latent interventions, and matched budget expansions, we find that current evidence most strongly supports H1 as a default working hypothesis rather than as a task-independent verdict. We therefore make two recommendations: the field should treat latent-state dynamics as the default object of study for LLM reasoning, and it should evaluate reasoning with designs that explicitly disentangle surface traces, latent states, and serial compute.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。