用因果审计发现大模型的思维链常是假装推理,实际靠隐藏路径做决策。
Bypassing the Rationale: Causal Auditing of Implicit Reasoning in Language Models
- 通过激活拼接技术,逐层检测思维链对决策的真实影响。
- 多数模型的思维链只在特定层起作用,且存在完全绕过现象。
- 专门训练推理的模型更可信,专家混合模型则呈现分布式推理特征。
思维链(CoT)提示被广泛视为增强模型推理能力并提供透明性的工具,但其行为提升并不意味着模型内部计算真正依赖生成的推理文本——模型可能产生流畅的推理过程,却将关键决策隐藏在潜空间路径中。本文提出一种基于激活拼接的因果、逐层审计方法,引入“思维链中介指数”(CMI),通过比较替换思维链词元隐藏状态与匹配控制块对性能的影响,量化其因果作用。在多个模型族(Phi、Qwen、DialoGPT)及不同规模下,我们发现思维链影响通常局限于狭窄的“推理窗口”,并识别出即使推理文本合理,CMI仍接近零的“绕过”情形。此外,显式训练推理的模型表现出更强且更结构化的中介作用,而混合专家模型则显示分布式的中介模式,符合路由式计算特性。结果表明,思维链的可信度在模型和任务间差异显著,无法仅凭行为表现推断,呼吁在将思维链作为透明性信号时进行因果、逐层审计。
原文摘要 · Abstract (English)
Chain-of-thought (CoT) prompting is widely used as a reasoning aid and is often treated as a transparency mechanism. Yet behavioral gains under CoT do not imply that the model's internal computation causally depends on the emitted reasoning text, i.e., models may produce fluent rationales while routing decision-critical computation through latent pathways. We introduce a causal, layerwise audit of CoT faithfulness based on activation patching. Our key metric, the CoT Mediation Index (CMI), isolates CoT-specific causal influence by comparing performance degradation from patching CoT-token hidden states against matched control patches. Across multiple model families (Phi, Qwen, DialoGPT) and scales, we find that CoT-specific influence is typically depth-localized into narrow "reasoning windows," and we identify bypass regimes where CMI is near-zero despite plausible CoT text. We further observe that models tuned explicitly for reasoning tend to exhibit stronger and more structured mediation than larger untuned counterparts, while Mixture-of-Experts models show more distributed mediation consistent with routing-based computation. Overall, our results show that CoT faithfulness varies substantially across models and tasks and cannot be inferred from behavior alone, motivating causal, layerwise audits when using CoT as a transparency signal.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。