用稀疏自编码器揭示大模型推理与直答的神经机制差异
Thinking vs. NoThinking: Towards Interpreting Reasoning Mechanisms of Large Language Models via Sparse Autoencoders

- 通过稀疏自编码器分析模型中间表示,对比推理与直答模式
- 发现推理依赖稀疏高强度特征,直答则偏向扩散式符号操作
- 揭示推理过程脆弱且依赖特征协同,适合研究模型可解释性
尽管采用思维链(CoT)的大语言模型展现出更强的推理能力,但其显式推理模式与直接作答模式之间的神经机制差异仍不明确。为解构这一认知过程,我们对 DeepSeek-R1-Distill-Qwen-7B 模型在三个不同难度层级的数学任务中的中间表示,应用 Top-K 稀疏自编码器(SAEs)。观察发现:推理模式依赖稀疏且高强度的特征激活,驱动独立于问题复杂度的语义推导;而直答模式则呈现适应性强、分布广的特征模式,更注重符号操作。因果干预实验表明:抑制最活跃的三个稀疏特征后,出现三类规律:(i) 推理与句法结构紧密耦合,干预导致 \LaTeX{} 和框式答案格式持续退化;(ii) 推理模式受扰后表现出补偿性过度生成,伴随元认知提示增多及重复低信息续写;(iii) 一致的思维链行为依赖专用特征间的脆弱协作,扰动下产生独特失效模式,但输出结构始终受损。
原文摘要 · Abstract (English)
While Large Language Models (LLMs) employing Chain-of-Thought (CoT) exhibit superior reasoning capabilities, the neural mechanisms distinguishing this explicit Thinking mode from direct answer generation (NoThinking mode) remain poorly understood. To deconstruct this cognitive process, we apply Top-K Sparse Autoencoders (SAEs) to the intermediate representations of DeepSeek-R1-Distill-Qwen-7B and examine the model's divergent behaviors across math-solving tasks of three distinct difficulty levels. Observationally, we identify a clear distinction in how the model functions under two reasoning modes: Thinking mode relies on sparse and high-intensity feature activations driving verbal deduction independent of problem complexity, whereas NoThinking mode exhibits an adaptive and diffuse pattern prioritizing symbolic manipulation. Causally, suppressing the three most active sparse features by Total Activation Volume reveals three principles: (i) reasoning and syntactic structure are tightly coupled, as interventions consistently degrade \LaTeX{} and boxed-solution formatting; (ii) Thinking responds to disruption with compensatory over-generation marked by increased metacognitive cues and repetitive, low-information continuations; and (iii) coherent CoT behavior depends on a fragile coordination among specialized features, yielding distinct failure modes under perturbation but a consistently impaired output structure.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。