揭示推理模型如何通过中间推理步骤影响最终答案生成。
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
- 通过实证实验验证显式推理提升答案质量。
- 注意力分析发现答案关注推理过程,尤其在中层有追踪推理轨迹的注意力头。
- 激活修补干预证明关键推理内容可直接改变最终答案。
大型推理模型(LRMs)在生成最终答案的同时会输出显式的推理链条,但这些推理过程对答案生成的影响程度尚不明确。本文针对三组蒸馏后的DeepSeek R1模型,开展三阶段研究:首先通过实证评估发现,包含显式推理能持续提升跨领域答案质量;其次注意力分析显示,答案词元显著关注推理词元,特定中层推理聚焦头(RFHs)紧密追踪推理轨迹,包括自我反思信号;最后通过激活修补的机制干预,证实对关键推理词元施加扰动可可靠改变最终答案,验证了从推理到答案的信息单向功能性流动。研究深化了对LRM如何利用中间推理生成输出的理解。数据与代码已公开于 https://aka.ms/R2A-code。
原文摘要 · Abstract (English)
Large Reasoning Models (LRMs) generate explicit reasoning traces alongside final answers, yet the extent to which these traces influence answer generation remains unclear. In this work, we conduct a three-stage investigation into the interplay between reasoning and answer generation in three distilled DeepSeek R1 models. First, through empirical evaluation, we demonstrate that including explicit reasoning consistently improves answer quality across diverse domains. Second, attention analysis reveals that answer tokens attend substantially to reasoning tokens, with certain mid-layer Reasoning-Focus Heads (RFHs) closely tracking the reasoning trajectory, including self-reflective cues. Third, we apply mechanistic interventions using activation patching to assess the dependence of answer tokens on reasoning activations. Our results show that perturbations to key reasoning tokens can reliably alter the final answers, confirming a directional and functional flow of information from reasoning to answer. These findings deepen our understanding of how LRMs leverage reasoning tokens for answer generation, highlighting the functional role of intermediate reasoning in shaping model outputs. Our data and code are publicly available at \href{https://aka.ms/R2A-code}{this URL}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。