arXiv:2608.22332cs.CL2026-08

通过逐次激活修补,揭示链式思维中关键注意力头的动态作用

Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching

论文配图:Mechanistic Interpretability of Chain-of-Thought Reasoning via Sequential Activation Patching
图 1 · 摘自论文原文
  • 设计序列激活修补框架,追踪推理过程中的注意力头变化
  • 发现多个注意力头协同支持推理轨迹维持与答案生成
  • 适用于研究大模型内部推理机制的学者与开发者

大型语言模型在链式思维(CoT)提示下展现出强大的问题求解能力,但其内部工作机制仍不清晰。本文研究了CoT相关因果效应在生成推理轨迹中的分布位置,以及哪些注意力头对最终答案计算有贡献。由于CoT推理跨越多个生成标记,传统单点静态激活修补无法刻画其时序特性。为此,我们提出一种序列激活修补框架,通过词性引导分析追踪注意力头在不同标记位置的激活,并聚合其影响。进一步引入序列多头修补方法,评估分布式头组的联合贡献,结合跨问题与随机激活对照。定向零消融实验表明,这些识别出的头在成功生成答案中具有功能性重要性,影响推理轨迹维持、答案锚定、示例-目标分离和数值生成等多个重叠机制。整体结果为与CoT条件计算相关的分布式推理支持子电路提供了证据。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate remarkable problem-solving capabilities when guided by Chain-of-Thought (CoT) prompting, yet the internal mechanisms underlying these improvements remain poorly understood. In this work, we investigate where CoT-related causal effects emerge across the generated reasoning trajectory and which attention heads carry signals that contribute to final-answer computation. Because CoT reasoning unfolds over multiple generated tokens, standard activation patching at a single static token position is insufficient to characterize these temporally distributed effects. To address this limitation, we introduce a sequential activation patching framework that traces CoT-conditioned attention-head activations across token positions and aggregates their effects using Part-of-Speech-guided analysis. We further introduce Sequential Multi-Head Patching to evaluate the joint contribution of distributed head sets, together with cross-question and random activation controls. Targeted zero-ablation experiments show that the identified heads are functionally important for successful answer generation and affect several overlapping mechanisms, including reasoning-trajectory maintenance, answer anchoring, exemplar-target separation, and numerical generation. Overall, our results provide evidence for distributed reasoning-support sub-circuits associated with CoT-conditioned computation.

链式思维可解释性注意力机制大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。